The frontier models get the headlines, but a lot of the value I see shipping is coming from the other direction: making models smaller. Distillation, quantization, and small task-specific models don't trend on social media, yet they're often what makes an AI feature actually viable in production.
Why smaller wins in production
The reasons are practical. A smaller model is cheaper per call, faster to respond, and can run closer to the user — sometimes on-device, which sidesteps a whole class of latency and data-privacy problems. When a task is narrow and well-defined, a compact model fine-tuned for it can match or beat a general giant, at a tiny fraction of the cost.
Quantization and distillation
Quantization is the least glamorous, highest-return trick in the box: represent weights in lower precision and you shrink memory and speed up inference, often with negligible quality loss. Distillation goes further — train a small "student" to mimic a large "teacher," capturing most of the behaviour in a package you can afford to run at scale.
Combine sizes, don't pick one
The strongest systems combine sizes rather than choosing one. Use retrieval to feed a small model the exact context it needs instead of paying a large model to memorise the world; keep a big model in reserve for the genuinely hard fraction of requests. Efficiency and capability aren't opposites — they're layers.
The real production question
It's worth internalising: the question that matters in production is rarely "how capable is the biggest model," but "how little compute can I spend and still clear the bar?" That's where a lot of the durable engineering is right now.
Sources & further reading
- Hinton, Vinyals & Dean, Distilling the Knowledge in a Neural Network (2015) — the original distillation paper.
- Sanh et al., DistilBERT (2019) — a distilled model at ~40% the size, retaining most of the behaviour.
- Dettmers et al., LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale (2022) — quantization without a quality cliff.
Editorial note — A trends / opinion piece on established techniques (quantization, distillation, retrieval, edge inference). No vendor-specific claims or numbers are asserted.


