Claude Fable 5.1 posted a record 66 on Artificial Analysis's Intelligence Index on September 1. At max effort it costs $3.76 per index task, 20 percent more than Fable 5, after a 75 percent cache-read cut that only saved about $1.40. The last Intelligence Index point is billed as extra thinking, not a cheaper token.
On September 1, Artificial Analysis published the first independent read of Anthropic’s Claude Fable 5.1. At maximum reasoning effort the model scored 66 on the Intelligence Index — four points above Fable 5’s 62, three above Anthropic’s own Claude Opus 5 at 63, and the highest number the firm has recorded. It leads the house’s agentic knowledge-work suites, posts the strongest SciCode mark AA has seen, and arrived with cache-read pricing cut from $1.00 to $0.25 per million tokens. List input and output stayed $10 and $50. That is the efficiency press release.
The invoice is less photogenic. Fable 5.1 at max effort costs $3.76 per Intelligence Index task, 20 percent above Fable 5’s $3.14 and about 1.6 times Opus 5’s $2.34. The cache cut saves roughly $1.40 on AA’s workload, concentrated where agents re-read a cached prefix. Without it the same task would have run about $5.16. Anthropic is not raising the meter. The model is running longer on it.
The Last Index Point Is an Output Purchase
Two token facts get collapsed into one slogan, and they are not the same. Against Fable 5, AA attributes the higher bill to about 1.7 times the output tokens — billed at $50 per million, the expensive line in the stack. Separately, Fable 5.1’s own five effort settings span an 11-fold range in output, from 13.1 million tokens at low effort (Index 58) to 143.7 million at max (Index 66). That 11x is how you climb the same model’s ladder, not how the successor compares to its parent.
The last rung is the expensive one. At xhigh effort Fable 5.1 scores 65 at $2.72 per task — a point below max, $1.04 cheaper. Buyers who treat “best ever measured” as a default setting are paying a dollar a task for a single Index point. Anthropic’s own docs already whisper the same arithmetic: start on Opus 5; reach for Fable 5.1 when Opus evals still fall short. Per-message effort, now in beta, exists because nobody can afford max on every turn.
This is not a uniquely Anthropic vice. Chain-of-thought and multi-step verification are token-expensive by design. China’s efficiency path still sits as the other pole: comparable capability without Western token girth. The U.S. frontier is still improving by thinking longer, then discounting the cheap line (cache) while the expensive line (output) grows.
Sticker tokens got cheaper. State-of-the-art answers did not.
OpenAI’s GPT-6 Astra, sold this week as computer-use rather than a deeper reasoner, makes the same effective move in a different currency. Standard API list is $10 per million input and $50 per million output — the same headline as Fable — against a prior GPT-5.6 Sol schedule that sat well below that pair. AA’s early Astra note is the mirror image of Fable: fewer tokens than Sol for similar Index performance, outweighed by higher prices. One lab got more verbose. The other got more expensive per token. Neither shipped a cheaper frontier.

Deflation Was the Capex Story. Routing Is the Bill.
The two-year investment narrative has been a deflation curve: intelligence per dollar falls, which is what makes hundreds of billions of data-center steel rational today. That thesis needs quality-adjusted cost per task to fall, not just the cache line on a price card. Fable 5.1 and Astra show sticker discounts arriving on schedule while the cost of a state-of-the-art answer holds or rises. That is the same gap sitting under thin consumer conversion and ghost racks and under a Street that can shrug at a record Nvidia quarter. Capex is a bet on future cheap intelligence. This week’s flagships say the cheapness is arriving as a routing problem, not as a falling bill at max effort.
The second-order effect is already in the product copy. Anthropic is selling Fable 5.1 for “read the evidence, weigh the considerations” work and pointing volume and execution at faster, cheaper models. Enterprises will run portfolios: judgment at high effort, clicks and drafts elsewhere. That is rational. It also means the savings, or the waste, show up one layer up — in the switchboard that decides which model, which effort, which cache — not in any single lab’s token sale. Packaging still rations the silicon. Verbosity now rations the invoice.
The audit is simple. Treat cache cuts as a discount on rereading, not on thinking. Treat max effort as a bid for the last Index point. If the deflation story is true, next quarter’s frontier should cost less per task at the same score, not merely print a smaller cache number beside a longer thought.
Continue reading
Sources
Artificial Analysis evaluation of Claude Fable 5.1 (Sept. 1, 2026); Anthropic Claude Fable 5.1 pricing and model docs; OpenAI GPT-6 Astra Standard API list prices (Sept. 3, 2026)