Every time a new model generation lands, the temptation is to reach for the biggest one and move on. The Claude 5 family — with frontier tiers alongside lighter, faster options like Fable 5 (and the compact Haiku line) — makes that instinct more expensive than it needs to be. The interesting engineering question isn't "which model is best"; it's "which model is right for this call."
Three buckets
I think about it as three buckets. A frontier tier for genuinely hard reasoning — architecture, ambiguous requirements, multi-step problem solving. A workhorse mid-tier that handles the majority of production traffic. And a fast, low-cost tier — this is where a model like Fable 5 shines — for the high-volume, latency-sensitive work: classification, routing, extraction, short drafting, and the first pass of an agent loop that a bigger model later checks.
The leverage is in routing
The leverage is in routing between them, not in standardising on one. A cheap model that answers in a fraction of the time can field 80% of requests, escalating only the hard ones. Done well, you get most of the quality at a fraction of the cost and latency — and your users feel the speed.
# Route by difficulty: cheap tier first, escalate only when needed.
def answer(request):
draft = fable5.run(request) # fast, low-cost first pass
if draft.confidence >= THRESHOLD:
return draft
return frontier.run(request) # escalate the hard fraction onlyEvaluation, not vibes
None of this replaces evaluation. Before I trust a smaller tier on a task, I want a small eval set that reflects real inputs, and a clear metric for "good enough." Model choice is a hypothesis; the eval is how you find out you were wrong cheaply.
A wider menu, not a bigger hammer
So treat a new generation less as a leaderboard and more as a wider menu. The win in the Claude 5 era isn't a single number going up — it's having a fast tier good enough to move real work off the expensive path.
Sources & further reading
- Anthropic, Models overview — the authoritative place for tiers, context windows, and pricing (which change; read them there, not here).
- Chen, Zaharia & Zou, FrugalGPT: How to Use LLMs While Reducing Cost and Improving Performance (2023) — a formal look at cascading cheap models into expensive ones.
Editorial note — This is a framing piece, not a spec sheet. I don't quote context windows, prices, or benchmark numbers here — those change and should be read from official documentation.


