A client wanted to run a 70B parameter model. Their budget: one consumer GPU.
Three years ago, I'd have said "impossible." Last month, I delivered it.
Quantization has changed the game completely. INT8 gets you most of the quality at half the memory. INT4 is surprisingly good for many tasks. There are even INT2 approaches that work for specific use cases.
The real magic is in the tools: llama.cpp runs quantized models on CPUs. GPTQ and AWQ handle GPU quantization elegantly. Ollama makes running local LLMs as simple as a single command. TensorRT squeezes every drop of performance from NVIDIA hardware.
Here's the counterintuitive finding: a well-quantized 7B model frequently beats an unoptimized 13B model. Because the 7B model fits entirely in GPU memory and avoids the performance cliff of memory swapping.
Size isn't everything in AI. A model that runs fast and cheap in production beats a model that requires a data center and a massive budget.
Efficiency is an engineering skill. And in 2026, it's one of the most valuable ones.