Not every query deserves a GPT-4 response. Most don't.
I built a routing system that sends simple queries to a tiny model (1.5B params), medium-complexity queries to a 7B model, and only routes genuinely complex queries to a frontier model.
The router itself is a small classifier trained on query complexity features: length, vocabulary diversity, number of entities, presence of reasoning indicators.
Result: 60% of queries handled by the tiny model. 30% by the mid-tier. Only 10% need the expensive model. Average response quality stayed the same. Cost dropped by 78%.
This pattern — LLM routing or cascading — is how every serious production system should work.
Think about it: using GPT-4 to answer "what time does the store close?" is like hiring a PhD to answer the phone. The information is in the FAQ. A simple retrieval or tiny model handles it perfectly.
Reserve your expensive compute for queries that actually need reasoning, nuance, or complex synthesis.
The implementation isn't complicated: a classifier for routing, a fallback mechanism (if the small model's confidence is low, escalate), and monitoring to adjust the routing thresholds over time.
Smart routing is the highest-ROI optimization in any LLM application. Build it early.