Opus 5 vs Fable 5
Opus 5 shipped yesterday at half of Fable 5's price and beat it on most benchmarks. Here is which one you should actually use.
Anthropic shipped Claude Opus 5 yesterday.
The pitch was one line: it comes close to Fable 5’s intelligence at half the price.
That sounds like normal launch talk. It is not. On most benchmarks the cheaper model actually wins.
So the obvious question is whether you should switch. Here is the full comparison, using Anthropic’s own published numbers, and a clear answer at the end.
Let’s go.
The Price
Opus 5 costs $5 per million input tokens and $25 per million output.
That is the same price as Opus 4.8, and exactly half of Fable 5 on input.
There is also a Fast mode at double the rate that runs about 2.5 times quicker. Opus 5 is now the default model on Claude Max and the strongest one on Claude Pro. It ships with a 1 million token context window.
So before we look at a single benchmark, the setup is simple. One model costs twice as much as the other.
Where Opus 5 Wins
Most places. This is the surprising part.
Agentic coding. On Frontier-Bench v0.1, Opus 5 scored 43.3%. Fable 5 scored 33.7%. Opus 4.8 scored 18.7%. So the new model more than doubled its own predecessor and beat the flagship above it.
Novel problem solving. On ARC-AGI-3, Opus 5 scored 30.2%. GPT-5.6 Sol scored 7.8%. Opus 4.8 scored 1.5%.
That test is worth explaining. ARC-AGI-3 drops a model into a small game world with no instructions and no stated goal. It has to work out the rules by playing.
When it launched in March 2026, every frontier model on Earth scored under 1%. Human testers solved all of it.
Four months later, one model is at 30%. That is not a small step.
Knowledge work. 1,861 Elo on GDPval-AA v2, against Fable 5’s 1,736.
Computer use. 70.6% on OSWorld 2.0, beating Fable’s best score at roughly a third of the cost.
Maths. Anthropic ran it on all six IMO 2026 problems with no tools and no agent setup. It scored 42 out of 42. That is gold medal level, well above the 29 out of 42 cutoff.
Anthropic’s official coding and knowledge work benchmark chart. This is the most useful image in the piece.
Anthropic’s ARC-AGI-3 chart. The 30.2 vs 7.8 vs 1.5 gap is the most dramatic visual of the launch.
Where Fable 5 Still Wins
This part got skipped in most launch coverage, so let me be clear about it.
Long autonomous work. Anthropic still recommends Fable 5 for the most advanced, long-horizon tasks. That is their own guidance, not mine.
One coding benchmark. On FrontierCode Main, Fable 5 edges it 53.5% to 53.4%. That is a coin flip, but worth noting that Anthropic lost four tests in their own published table and printed all of them anyway. I respect that.
Specialist domains. Fable and Mythos still lead on legal, health, and frontier cybersecurity work.
Opus 5 is deliberately built with lighter cyber safeguards so developers can find and fix vulnerabilities, and Anthropic says plainly that it stays well behind Mythos 5 at developing exploits.
Finishing long tasks. Early testers reported a consistent pattern. On shorter tasks Opus 5 matched or beat Fable. On longer ones it was less ambitious and did not always deliver as complete a piece of work.
Anthropic’s cost per task chart. This is the one that carries the actual argument, that the cheaper model does more work per rupee.
The Thing Nobody Expected
Here is the strangest part of launch week.
The early reviews contradicted each other. Some testers called it a huge leap. Others said it argued with instructions, stopped before finishing, and did not work well with the skills and workflows they had built for earlier models.
Same model. Same week. Opposite verdicts.
The explanation showed up in the reviews themselves. The team at Every, which had a full week with it before launch, deleted all their existing skills and started from scratch.
The model immediately got much better. They also found that lower thinking levels worked better than higher ones, which is the opposite of what you would expect.
And on the same day, Anthropic’s own Claude Code team published that they had removed about 80% of their system prompt for these new models.
Eighty percent. Deleted. By the people who wrote it.
So here is the simple version.
Most of our prompting habits exist to fix old weaknesses. Spell out every step, because models used to skip them. Add “check your work”, because they used to lie. Break everything into tiny tasks, because they used to lose track.
Opus 5 does not need most of that. And when you feed it those instructions anyway, you are not helping. You are getting in the way.
That is why the same model looks brilliant to someone starting fresh and broken to someone with two years of carefully tuned workflows.
Those reviews were not really measuring the model. They were measuring how much old scaffolding each person was still carrying.
Which One Should You Use
Straight answers.
Use Opus 5 for most things. Daily coding, agent work, computer use, knowledge work, anything inside a large repo. It wins on the benchmarks and costs half as much. This is your new default.
Use Fable 5 for long autonomous runs and specialist work. If you are handing over a task that runs for hours without you, or you need legal, health, or deep security work, Fable is still ahead.
Try low effort before high. This is the least obvious lesson from launch week. Start at medium on the effort toggle. Go high only when the problem genuinely needs it. More thinking is not automatically better here.
Delete your prompts before you judge it. Run one real task with no skills, no long system prompt, no step by step babysitting. That is your true baseline. Judge the model on that, not on scaffolding you built for a different one.
Watch token burn, not token price. Several testers found this model hungry. Half price per token means nothing if it uses three times as many. Measure cost per finished task, not cost per token.
One last thing.
The interesting part of this launch was not that a cheaper model beat an expensive one. That will keep happening, and fast.
It was watching experienced people test a better tool with old habits, and conclude the tool was worse.
The models are improving faster than our habits are. So when the next one lands, and it will land soon, run it clean first. No skills, no scaffolding, nothing you built for the last one.
Then decide.
Manav
P.S. If you tried Opus 5 and did not like it, reply and tell me what was in your system prompt. I will bet the prompt was the problem. New here? Subscribe to Tensor Protocol.
Sources, check every number yourself:
Anthropic’s official Opus 5 announcement. Pricing, benchmarks, alignment audit, cyber safeguards. Primary source for everything above.
VentureBeat on the launch. Independent coverage of the Frontier-Bench and OSWorld numbers.
MarkTechPost technical breakdown. Where the IMO 42/42 result and the verified ARC-AGI-3 score come from.
Every: Vibe Check on Opus 5. The most critical review of launch week, and the one that found the fix.
Artificial Analysis: Claude Opus 5. Independent cost per task data, updated continuously.





