← Back to library
ToolsOpen5 min read

Claude's Cache Pricing Just Dropped 75%: The 15-Minute Question to Ask Your AI Vendor

Anthropic cut prompt-caching costs 75% with the Claude Fable 5.1 release. If you run any AI tool through an API instead of a $20 chat app, this changes your bill automatically, or it should.


What happened: Anthropic shipped Claude Fable 5.1 on September 1, 2026, and cut prompt-caching read pricing by 75%. A cache read happens when an AI call reuses a chunk of context you already sent it, a system prompt, a Context Pack, a standard operating procedure, instead of processing it fresh every single time.

Why it matters for you: if your business runs any AI workflow through an API key instead of typing into Claude.ai or ChatGPT by hand, this is the gap between paying full price on every call and paying a fraction of it on every repeat call. Support macros, weekly report generators, content pipelines, a chatbot fed your business's FAQ, anything that sends the same block of context over and over now costs a lot less to run at scale.

Who this doesn't touch: if you and your team just type prompts into a chat app, this happens behind the scenes and there's nothing for you to configure. It only matters for automations built on the API.

The 15-minute move: if you've paid a developer, or used a platform like Make, Zapier, or n8n to build a custom AI tool for your business, ask one question: does our setup use prompt caching, and did the Fable 5.1 pricing cut apply automatically? If whoever built it doesn't know, that's the tell your integration was built without cost controls in mind. Worth a second look before your next invoice, especially if you've got a tool that pastes the same long context (a brand voice doc, an SOP, a product catalog) into every single call.

Sources: Local AI Zone, September 2026 AI Model Updates, AI Release Tracker