The New AI Model Tiers Changed the Math: A 15-Minute Pass to Route Every Task to the Right One
In about three weeks OpenAI, Anthropic, and Google all shipped new model tiers at very different prices and speeds. Here is the simple routing rule that puts boring work on the cheap fast model and saves the expensive one for the calls that matter.
What shipped. In about three weeks, all three major labs pushed new model tiers. OpenAI released the GPT-5.6 family on July 9: Sol at the top, Terra in the middle, and Luna as a low-cost, fast tier. Anthropic shipped Claude Sonnet 5 as its scalable default and Claude Fable 5 as the higher-capability option. Google put out Gemini 3.5 Flash with a Pro version on the way.
The so-what for operators. The gap between the cheapest tier and the flagship is now large, in both price and speed. The cheap, fast tiers are good enough for the boring 80% of what you actually run: cleaning up notes, drafting replies, summarizing a call, formatting a list. Paying flagship rates for that work is money and time walking out the door. The flagship earns its keep on the hard 20%: a pricing decision, a contract read, a message going to your best client.
The move (15 minutes). Write a one-page routing rule for your team:
- Fast, cheap tier (Luna, Sonnet 5, Gemini Flash): first drafts, summaries, reformatting, brainstorming, internal notes.
- Flagship tier (Sol, Fable 5, Gemini Pro): anything a customer sees unedited, any number you will act on, any legal or financial read, any final copy.
- Always cross-check high-stakes output with a second model before it ships.
Pin that list where your team works. Most tools now have a model picker right in the chat box, so switching is one click. You keep flagship quality where it counts and stop overpaying everywhere else.
Keep client data in your private or business AI workspace, whichever tier you pick.
Routing rule posted to the team wiki, 8/29:
Cheap tier (Luna / Sonnet 5 / Gemini Flash): call summaries, meeting notes cleanup, first-draft SOPs, reformatting spreadsheets into lists, internal Slack replies, brainstorm rounds.
Flagship tier (Sol / Fable 5 / Gemini Pro): LOI and contract review, any rehab budget or rent number going into an underwriting model, investor updates, final website and ad copy, anything a client reads unedited.
Cross-check rule: before a flagship output ships, paste it into a second lab's flagship and ask "what is wrong here." Two minutes, catches the confident-and-wrong ones.
Sample month at a 6-person shop: roughly 1,100 AI tasks. About 900 were reformatting or drafting. Moving those off the flagship cut the tier's usage by two-thirds.
Owner note: the picker is in the chat box. One click. No new tool.