← Back to library
ToolsOpen5 min read

The New AI Model Tiers Changed the Math: A 15-Minute Pass to Route Every Task to the Right One

In about three weeks OpenAI, Anthropic, and Google all shipped new model tiers at very different prices and speeds. Here is the simple routing rule that puts boring work on the cheap fast model and saves the expensive one for the calls that matter.


What shipped. In about three weeks, all three major labs pushed new model tiers. OpenAI released the GPT-5.6 family on July 9: Sol at the top, Terra in the middle, and Luna as a low-cost, fast tier. Anthropic shipped Claude Sonnet 5 as its scalable default and Claude Fable 5 as the higher-capability option. Google put out Gemini 3.5 Flash with a Pro version on the way.

The so-what for operators. The gap between the cheapest tier and the flagship is now large, in both price and speed. The cheap, fast tiers are good enough for the boring 80% of what you actually run: cleaning up notes, drafting replies, summarizing a call, formatting a list. Paying flagship rates for that work is money and time walking out the door. The flagship earns its keep on the hard 20%: a pricing decision, a contract read, a message going to your best client.

The move (15 minutes). Write a one-page routing rule for your team:

  • Fast, cheap tier (Luna, Sonnet 5, Gemini Flash): first drafts, summaries, reformatting, brainstorming, internal notes.
  • Flagship tier (Sol, Fable 5, Gemini Pro): anything a customer sees unedited, any number you will act on, any legal or financial read, any final copy.
  • Always cross-check high-stakes output with a second model before it ships.

Pin that list where your team works. Most tools now have a model picker right in the chat box, so switching is one click. You keep flagship quality where it counts and stop overpaying everywhere else.

Keep client data in your private or business AI workspace, whichever tier you pick.