All posts

Opus 5 beats Fable 5 on coding benchmarks at half the cost

Manaal KhanJuly 25, 2026 at 1:32 AM5 min read
Opus 5 beats Fable 5 on coding benchmarks at half the cost

Key Takeaways

Opus 5 beats Fable 5 on coding benchmarks at half the cost
Source: The Decoder
  • Opus 5 scores 43.3% on Frontier-Bench agentic coding, beating Fable 5 (33.7%) and GPT-5.6 Sol (34.4%)
  • Token pricing stays at $5/$25 per million input/output tokens, half of Fable 5's rates
  • On ARC-AGI-3 novel problem-solving, Opus 5 hits 30.2%, nearly 4x higher than GPT-5.6 Sol's 7.8%

Anthropic's new flagship model Claude Opus 5 topped Fable 5 in agentic coding and knowledge work benchmarks while keeping token prices at half the rate. The model scores 43.3% on Frontier-Bench v0.1 terminal coding, compared to Fable 5's 33.7% and GPT-5.6 Sol's 34.4%. It becomes the default on Claude Max and the most capable model on Claude Pro.

The release responds to pricing pressure from OpenAI's GPT-5.6 Sol and Chinese competitors. Anthropic is betting that near-Fable performance at Opus pricing will shift usage patterns for teams running high-volume inference.

Advertisements

How does Opus 5 pricing compare to Fable 5?

Opus 5 costs $5 per million input tokens and $25 per million output tokens. That's unchanged from Opus 4.8. Fable 5 runs double: $10 input, $50 output. A new Fast Mode boosts speed by 2.5x but doubles the price.

ModelInput (per MTok)Cache Writes 5 minCache Writes 1 hrCache HitsOutput (per MTok)
Claude Fable 5$10$12.50$20$1$50
Claude Mythos 5$10$12.50$20$1$50
Claude Opus 5$5$6.25$10$0.50$25

Token rates don't tell the full story. Opus 4.7 ended up costing 30 to 40 percent more per task than Opus 4.6 despite identical base rates. The reason: token efficiency varies by model. A model that reasons longer or outputs more tokens per answer eats into its pricing advantage.

Image (Source: The Decoder)
Image (Source: The Decoder)

Effort settings: higher isn't always better

Anthropic offers five effort settings: low, medium, high, xhigh, and max. The company says Opus 5 delivers better value than its predecessor at every level. But here's the catch: on two benchmarks, max effort scored worse than xhigh despite costing more.

The drop shows up on Frontier-Bench v0.1 and the Artificial Analysis Coding Agent Index. Anthropic's prompting guide recommends starting with low or medium for most tasks, reserving xhigh for coding and agentic work. Max effort, it turns out, can burn tokens without improving results.

Image (Source: The Decoder)
Image (Source: The Decoder)
Also Read
Anthropic's Claude Cookbook: 15 agent patterns for startups

Practical agent patterns to pair with Opus 5's agentic capabilities

Where Opus 5 leads and where it trails

Opus 5 doesn't win everywhere. On DeepSWE v1.1 agentic coding, GPT-5.6 Sol leads with 72.7%, followed by Fable 5 at 69.7% and Opus 5 at 68.8%. Fable 5 outperforms on health tasks; Mythos 5 beats it on legal benchmarks.

On knowledge work (GDPval-AA v2), Opus 5 leads with an Elo of 1,861. Fable 5 trails at 1,747, GPT-5.6 Sol at 1,736. For teams running document analysis, summarization, or research workflows, that gap matters.

Image (Source: The Decoder)
Image (Source: The Decoder)
Advertisements

The ARC-AGI-3 outlier

The biggest surprise in Anthropic's benchmarks is ARC-AGI-3, which measures novel problem-solving without memorized patterns. Opus 5 scores 30.2%. Opus 4.8 managed 1.5%. GPT-5.6 Sol hit 7.8%. That's nearly a 4x gap over the next-best model.

There's no Fable 5 result on this benchmark, and whether a large lead on a synthetic test translates to real-world use remains an open question. Still, a jump from 1.5% to 30.2% between model versions suggests something changed in how Anthropic trains for generalization.

Image (Source: The Decoder)
Image (Source: The Decoder)

Cybersecurity: deliberately limited

Opus 5 falls behind Mythos 5 on cybersecurity tasks. Anthropic says it deliberately avoided training Opus 5 on cyber, same as its predecessor. The model finds vulnerabilities about as well as Mythos 5 but performs much worse at exploit development.

This is a policy choice, not a capability gap. Anthropic is trading offensive security performance for reduced misuse risk. Teams needing exploit capabilities will have to look elsewhere or use Mythos 5 under limited availability.

Image (Source: The Decoder)
Image (Source: The Decoder)

New capabilities: tool building and visual analysis

Anthropic says Opus 5 can check and improve its own work through iteration, and build its own tools through code when it needs them. The model also shows improved performance on visual outputs and analyzing charts, diagrams, and other visual content.

For agentic workflows, tool building matters. A model that can write and execute helper functions mid-task reduces the need for pre-built tool libraries. Whether this works reliably in production is something early adopters will have to test.

Also Read
Claude Opus 5 matches Fable 5 performance at half the cost

More detail on Opus 5's positioning against Fable 5

ℹ️

Logicity's Take

For AI product teams, Opus 5 looks like the new default for coding agents and knowledge work. The math is simple: near-Fable performance at Opus pricing cuts inference costs roughly in half for those workloads. But don't set effort to max and walk away. The benchmark data shows xhigh often beats max at lower cost. Teams running high-volume inference should A/B test effort levels against their specific tasks. The real competition here is GPT-5.6 Sol, which leads on DeepSWE and comes at aggressive pricing. Opus 5's advantage is strongest on novel problem-solving and terminal coding.

Frequently Asked Questions

How much does Claude Opus 5 cost compared to Fable 5?

Opus 5 costs $5 per million input tokens and $25 per million output tokens. Fable 5 costs $10 input and $50 output, making Opus 5 half the price.

Does Opus 5 beat GPT-5.6 Sol on coding benchmarks?

On Frontier-Bench agentic coding, Opus 5 leads with 43.3% vs 34.4%. But on DeepSWE, GPT-5.6 Sol leads with 72.7% vs 68.8%.

What is the context window for Claude Opus 5?

Opus 5 keeps the same 1 million token context window as its predecessor.

Should I use max effort setting on Opus 5?

Not necessarily. Anthropic's benchmarks show max effort scores worse than xhigh on two tests despite costing more. Start with xhigh for coding tasks.

Can Opus 5 build its own tools?

Yes. Anthropic says Opus 5 can write and execute helper functions through code when it needs them during task execution.

ℹ️

Need Help Implementing This?

Want to integrate Claude Opus 5 into your product or compare it against GPT-5.6 Sol for your use case? Contact Logicity's AI implementation team for benchmarking support and deployment guidance.

Source: The Decoder / Matthias Bastian

M

Manaal Khan

Tech & Innovation Writer

Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.