AI Briefing: July 28, 2026 — Claude Opus 5 Outscores Fable 5 on Most Public Benchmarks. It Costs Half as Much.

A SECOND FLAGSHIP, HALF THE PRICE

Claude Opus 5 went live on July 24 across Claude.ai, where it's now the strongest model available to Pro subscribers and the new default for Max, alongside the API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry — one of Anthropic's broadest same-day rollouts to date. It runs with thinking on by default, carries a 1-million-token context window with up to 128K tokens of output, and is priced at $5 per million input tokens and $25 per million output tokens, identical to Opus 4.8's rate card and exactly half of what Fable 5 costs at $10/$50. Anthropic's own description leans on that gap directly: Opus 5 is built to land close to Fable 5's frontier intelligence without Fable 5's price. That's a notable thing for a company to say about its own top-tier product four weeks after launching it, and it's the frame the rest of this week's reaction has argued about.

THE BENCHMARK TABLE, IN FULL

The headline numbers hold up under a wider read, not just a cherry-picked one. On SWE-bench Verified, Opus 5's 97.0% edges out GPT-5.6 Sol's 96.2% and Fable 5's 95.0%. On the harder SWE-bench Pro, Opus 5 posts 79.2%, a fraction behind Fable 5's 80.3% but 14.6 points ahead of Sol's 64.6%. On Frontier-Bench, Anthropic's agentic-coding evaluation, it's 43.3% against Sol's 34.4% and Fable 5's 33.7%. On ARC-AGI-3, a benchmark built to resist memorization, Opus 5 scores 30.2% against Sol's 7.8% and Fable 5's 1.5% — roughly four times Sol's result and twenty times Fable 5's. On GDPval-AA v2, a knowledge-work evaluation, it's 1,861 against Fable 5's 1,747 and Sol's 1,736. Tallied across the twelve benchmarks Anthropic published at launch, independent trackers count Opus 5 ahead on nine of them, essentially tied on the rest.

EFFORT AS A DIAL, NOT A MODEL CHOICE

The feature Anthropic is pairing with those numbers is a per-request effort setting — min, low, medium, high, and max — that controls how much internal reasoning the model does before it answers, rather than forcing a choice between separate model names. A routine call can run on low for speed and cost; a hard debugging session can be pushed to max for the deepest reasoning the model has, with the choice persisting across a session and switchable mid-conversation in Claude Code via /effort. Anthropic also shipped a separate fast mode, billed at roughly double the standard rate, around $10/$50 per million tokens, that trades money for about 2.5x the response speed on tight interactive loops. Together, the two levers turn what used to be a decision about which Claude model to buy into a decision about how much of one model to spend on any given task — a more granular version of the tradeoff Fable 5's own credits system tried, and struggled, to manage.

THE HACKER NEWS PUSHBACK

The launch thread on Hacker News pulled in 1,378 points and 746 comments within hours, and a meaningful share of that volume was aimed at the benchmark table itself rather than the model. Commenters flagged that Opus 5 scored better on FrontierCode at medium effort than at high effort, even though more effort improved every other eval — a result that reads as either a genuine task-specific tradeoff or evaluation noise Anthropic hasn't fully explained. Others pointed to Anthropic's ECI capability index showing only a one-point gain over Opus 4.8, which several users called "incredibly underrated" given how much stronger the model feels in day-to-day use, and a handful accused Anthropic of inconsistent bolding in its own comparison tables to make Opus 5's wins look more uniform than they are. None of it overturns the headline scores. It does mean the twenty-times gap over Fable 5 on ARC-AGI-3 deserves the same skepticism this site has applied to every other single-benchmark claim this month, including Alibaba's on July 21 and Moonshot's on July 27 — a big number from a model's own vendor is a reason to test it yourself, not to repeat it.

WHAT THIS DOES TO FABLE 5'S PRICE TAG

This is the second time in three weeks Anthropic's own decisions have undercut Fable 5's premium positioning. On July 7, Fable 5 moved to credits-only pricing for every Claude subscriber, pulling back the promotional 50% weekly-usage inclusion it launched with less than a week earlier. Now, on numbers Anthropic itself published, a model priced at half of Fable 5's rate beats it on SWE-bench Verified, Frontier-Bench, and ARC-AGI-3, and trails it by a single point on SWE-bench Pro. Anthropic's official line is that Fable 5 still leads on the hardest, most open-ended agentic work Opus 5 hasn't been tested against in the same detail. That may hold up. But a company doesn't usually publish a comparison table this close between its $10-per-million and $5-per-million products unless it's decided the $10 tier needs to justify itself on more than a leaderboard — and for most teams choosing between the two, the leaderboard is now the wrong place to look.

THE SAFETY CARD IMPROVED, TOO

Away from the pricing story, Opus 5's system card shows real gains on the metrics Anthropic tracks for misuse resistance. On its automated behavioral audit, Opus 5 scored 2.30 for overall misaligned behavior, the lowest, meaning best, of any recent Claude release, ahead of Opus 4.8, Sonnet 5, and Fable 5 alike. On indirect prompt injection, the rate at which an attacker succeeds within fifteen attempts fell from 5.5% under Opus 4.8 to 2.0% under Opus 5, and in computer-use environments specifically, the attack success rate dropped from 7.14% to 0.54% when extended thinking is enabled. None of that made the Hacker News front page the way the benchmark dispute did, but it's the more durable claim of the two: it's a comparison against Anthropic's own prior models rather than a rival's, and it lines up with the general direction every major lab has reported on injection resistance this year.

WHAT THIS MEANS FOR TEAMS BUILDING ON AI

If Fable 5 was your default for anything short of the hardest agentic workloads, Opus 5 at half the price is worth a direct A/B on your own tasks this week, not a decision made from Anthropic's table alone — the FrontierCode effort anomaly and the ECI gap are real enough that your workload may not reproduce the twenty-times ARC-AGI-3 gap or the SWE-bench win outright. Budget the effort dial deliberately: default routine calls to low or medium, reserve high and max for the reasoning-heavy steps where they actually change the output, and treat fast mode as a latency tool for interactive sessions rather than a default. And if you're weighing Fable 5's $10/$50 tier against Opus 5's $5/$25 one for a new project, start the evaluation assuming they're close until your own numbers say otherwise — that assumption would have saved a lot of benchmark-table arguing this week.