Controversial take: most production systems don't need GPT-4 class models.
I replaced a GPT-4 API call with a fine-tuned 3B parameter model for a client's customer categorization task. Same accuracy. Cost went from $400/day to $12/day. Latency dropped by 80%.
Small Language Models (1B-7B parameters) are having a moment for good reason. Phi-3, Gemma 2, Qwen 2.5, Llama 3.2 — these aren't "lesser" models. They're right-sized for specific tasks.
The math is simple: if your use case is well-defined (classification, extraction, formatting, domain-specific Q&A), a well-tuned small model often matches a general-purpose giant model. But it runs on consumer hardware, costs pennies to operate, and responds in milliseconds.
Smart engineering isn't about using the biggest model. It's about using the right-sized model.
Start with the smallest model that could possibly work. Scale up only if it doesn't. This saves money, improves latency, and often forces you to think more carefully about your problem definition — which leads to better solutions anyway.
The best model isn't the biggest. It's the one that solves your problem at the lowest cost.