Stylized number 5 made of orange ribbons and gears – cover for Claude Sonnet 5

    Claude Sonnet 5: Agentic AI Goes Mainstream

    30. Juni 2026Updated: August 3, 20269 min readDeep Dive
    Till Freitag

    TL;DR:Claude Sonnet 5 (June 30, 2026) brings near-Opus-4.8 agentic performance at Sonnet pricing: $2/$10 per M tokens (intro until Aug 31), then $3/$15. Four weeks later Claude Opus 5 (July 24, $5/$25) lands on top: SOTA on Frontier-Bench, 3× ARC-AGI-3, best agent on Zapier AutomationBench – but slow, verbose and with a higher hallucination rate. Sonnet 5 stays the volume default, Opus 5 becomes the judgment layer."

    Till Freitag

    What Anthropic shipped today

    On June 30, 2026, Anthropic released Claude Sonnet 5 – the most agentic Sonnet model yet. Performance close to Opus 4.8, priced at Sonnet levels, available everywhere on day one: Free, Pro, Max, Team, Enterprise, Claude Code, and the API as claude-sonnet-5.

    ModelInput / Output (per 1M tokens)Availability
    Sonnet 5 (intro until Aug 31, 2026)$2 / $10Free, Pro, Max, Team, Enterprise, Claude Code, API
    Sonnet 5 (standard from Sep 1, 2026)$3 / $15same
    Opus 4.8 (reference)$5 / $25Premium

    Why this is a turning point

    The Sonnet 3.x generation was the on-ramp to agentic coding and tool use in 2024/25. After that, the biggest agentic jumps mostly happened in the Opus class – with Opus-class pricing to match.

    Sonnet 5 closes that gap. Anthropic itself frames it as "close to Opus 4.8 performance, at lower prices." In practice: workloads that previously required Opus (long coding sessions, computer use, multi-step agents) are now economical on Sonnet.

    This isn't only a price update – it's a default shift. Sonnet is now the default model in Free and Pro plans. Millions of users get agentic behavior with no opt-in.

    The benchmarks (per Anthropic)

    Anthropic compares Sonnet 5 against Sonnet 4.6 and Opus 4.8 in the launch post:

    • BrowseComp (agentic search): Sonnet 5 strictly improves over 4.6 and approaches Opus 4.8 at higher effort levels.
    • OSWorld-Verified (computer use): same pattern – Sonnet 4.6 trailed Opus 4.8, Sonnet 5 and Opus 4.8 now cover one range at different price/effort points.
    • Humanity's Last Exam, SWE-bench, tool use: per the system card, jumps across the board vs. 4.6, plus lower hallucination and sycophancy rates.

    If you want methodology context, we wrote about how AI benchmarks actually work (and where they break) – including Goodhart's Law and live arenas.

    What early-access partners report

    The launch post quotes Lovable, ClickHouse, Pace, Eve, and others. Common thread:

    • "Same output, fewer steps." Sonnet 5 reaches comparable results with fewer iterations – which matters because token consumption, not raw model quality, drives real cost.
    • "Stays on plan." On multi-step workflows (Salesforce updates + outreach, PR pipelines, insurance FNOL) Sonnet 5 finishes end-to-end more often instead of stalling.
    • "Checks its own output." Self-verification without explicit prompting – exactly the behavior people used to coax out with "please double-check your answer" reflection tricks (see our notes on Lovable's model routing).

    What changes for builders

    Three concrete consequences if you ship on Claude today:

    1. Switch your default model in Claude Code. For most coding sessions, Sonnet 5 is enough. Keep Opus 4.8 for the hardest brownfield tasks and anything that needs the top of the accuracy curve.
    2. Effort levels, not model switching. Sonnet 5 supports multiple effort levels (up to xhigh). Instead of ping-ponging between Sonnet and Opus, stay on one model and dial effort.
    3. Take the tokenizer note seriously. Sonnet 5 uses a new tokenizer; the same input can map to roughly 1.0–1.35× more tokens. Intro pricing is set so the transition is roughly cost-neutral – re-evaluate your pipelines once standard pricing kicks in.

    Update, August 3, 2026: Opus 5 lands four weeks later

    On July 24, 2026 Anthropic shipped Claude Opus 5 – barely four weeks after Sonnet 5. "Sonnet 5 is close to Opus 4.8" didn't become wrong, but the reference point moved. Opus 5 costs exactly what Opus 4.8 cost ($5 / $25 per M tokens) while, per Anthropic, coming "close to the frontier intelligence of Claude Fable 5 at half the price." It's the new default on Claude Max.

    ModelInput / Output (per 1M tokens)Role since August 2026
    Sonnet 5$3 / $15 (intro $2/$10 until Aug 31)Volume default, Free/Pro
    Opus 5$5 / $25Default on Max, strongest model on Pro
    Fable 5~2–3× Opus 5 cost per taskFrontier ceiling
    Mythos 5Premiumstill ahead on cyber/bio tasks

    The numbers that are actually new

    From the system card and the Artificial Analysis write-up:

    BenchmarkOpus 5Opus 4.8Fable 5GPT-5.6 Sol
    Frontier-Bench v0.143.321.133.834.4
    SWE-bench Multimodal59.438.454.1
    SWE-bench Pro79.269.28064.6
    OSWorld 2.070.655.766.162.6
    Zapier AutomationBench26.017.017.418.1
    ARC-AGI-330.21.57.8
    GDPval-AA v2 (Elo)1861159317471736

    Two things stand out. First, the jump is unevenly distributed: Opus 5 loses SWE-bench Pro to Fable 5 and ARC-AGI-2 to GPT-5.6 Sol, while doubling Opus 4.8 on Frontier-Bench and going 20× on ARC-AGI-3. Second, at the 2026 IMO (July 15–16) Opus 5 scored a perfect 42/42 per the system card – no agent harness, no tools. Gold cutoff was 29.

    The real difference vs. earlier models: judgment, not output

    Anyone who lived through 4.5 → 4.6 → 4.7 → 4.8 knows the pattern: higher scores, same behavior. Opus 5 breaks it. The thread running through Anthropic's early-access quotes and the independent reviews isn't "writes better code," it's "decides better when to stop."

    • Self-verification by default. A frontend team reports Opus 5 opened its own pages in a browser at desktop and phone widths, caught a product hidden below the mobile fold and an off-screen checkout button, and fixed both before handing the work back. Sonnet 5 verifies; Opus 5 builds its own test harness to verify.
    • Pushback instead of sycophancy. One tester reports Opus 5 rejected his proposed design in a rearchitecting session – and didn't fold when pressed. It narrowed the objection to a single design question and proposed a compromise. Exactly the property whose absence we described in Lovable's model routing.
    • Fewer turns for the same result. A financial-modeling team measured 9 percentage points higher accuracy with a third fewer turns and tool calls and 60% less time. A trading firm reports roughly a seventh of the reasoning tokens versus Opus 4.8.
    • Lovable itself reports +22% over Opus 4.7 on its hardest agentic coding tasks – and emphasizes reduced run-to-run variance above all. For a product that has to ship reproducible builds, consistency beats peak score.
    • Prompt-injection resistance. Indirect prompt-injection attack success drops from 5.5% to 2.0% (GPT-5.6 Sol: 20.0%); browser-use attacks from 31.5% to 3.7%. For anything giving agents access to real customer data, that's the most relevant number in the release.

    What experts criticize

    Launch week split in an unusual way – not over quality, but over ergonomics.

    Claire Vo put it bluntly: "Opus 5 is here… and I hate working with it. And yet in a blind taste test, I ranked it above every other model." Dan Shipper (Every) landed in the same place after a week of pre-release testing – "it's a hard model to love": it argued with instructions and stopped before the work was finished. Kieran Klaassen reports it broke his Compound Engineering setup, with pragmatic advice attached: drop the big mega-prompts and skill stacks, work at medium effort.

    Three hard numbers to know:

    1. Hallucination rate goes up. On AA-Omniscience, Opus 5 improves accuracy by 7 points over Opus 4.8 – while its hallucination rate rises 14 points to 50%. Anthropic's system card records the same trade at smaller scale (+11% accuracy, +6% hallucination). The model answers more often when it's uncertain. On coding with tests that's cheap; on customer communication it's the whole failure mode.
    2. It is slow, in the worst place. 52.6 output tokens/second (rank 117 of 190, class median 76) and – decisively – 68 seconds time-to-first-token at max effort against a class median of 2.81. That disqualifies max-effort Opus 5 from real-time UI outright. Fast mode buys ~2.5× speed at twice the base price.
    3. max is rarely the right setting. On AA-Briefcase, max effort averaged 36.2 minutes and 103 turns per task (Opus 4.8: 24.1 minutes, 55 turns) at $2.03 per index task – more than Opus 4.8 ($1.80) and Sonnet 5 ($1.53). Anthropic has since revised its own guidance to start at high.

    One methodological note most coverage skips: Anthropic ran its Frontier-Bench numbers with Opus 4.8 as a fallback on safety-classifier refusals, which caught roughly 5% of calls. A small slice of every "Opus 5" score was produced by Opus 4.8. We wrote about how details like this distort leaderboards in how AI benchmarks actually work.

    Migration: two things genuinely break

    Anthropic calls Opus 5 a drop-in replacement for Opus 4.8. Two exceptions:

    • Adaptive thinking is on by default. Thinking tokens now count against max_tokens – requests that previously had none can hit the limit.
    • Disabling thinking is capped. thinking: {"type": "disabled"} returns a 400 at xhigh or max effort. Opus 4.8 still accepted that combination.

    Web fetch and Priority Tier are also missing on Opus 5. New in beta: mid-conversation tool changes (swap tools without invalidating the prompt cache – relevant for subagent patterns) and automatic server-side fallbacks.

    Our routing rule since August

    • Sonnet 5 – volume: content, classification, routine refactors, PR reviews, anything with real-time expectations.
    • Opus 5 at high – long agent runs, architecture decisions, ambiguous tasks, anything where a wrong answer is expensive to unwind.
    • Opus 5 at max – only when runtime doesn't matter and the task genuinely sits at the ceiling.
    • Fable 5 / Mythos 5 – frontier remainder: cyber analysis, bio research, Project Glasswing.

    The honest one-liner for this release isn't "best coding model." It's: Opus 5 is the first model where the delta shows up primarily in judgment rather than output. That's exactly the point where agents shift from "does work" to "takes responsibility" – and the reason the hallucination number still deserves attention.

    How it fits the AI race

    Sonnet 5 was Anthropic's second major move in three weeks – after Claude Fable 5 & Mythos 5. With Opus 5 on July 24, that's three releases in four weeks and a very clear lineup:

    • Sonnet 5 – default agent for the broad market.
    • Opus 5 – judgment, long horizons, agentic responsibility.
    • Fable 5 / Mythos 5 – frontier, cyber and Project Glasswing use cases.

    We've added both releases to our timeline: The AI race in 43+ milestones.

    What we'll be watching

    • Real cost curves after the tokenizer change, especially for long agent runs.
    • Behavior in Claude Code over very long sessions – does self-verification hold up over hours?
    • Downstream effect of the Free/Pro default shift on the ecosystem, including Lovable, which is likely to route many flows through Sonnet 5 by default.

    Sources

    TeilenLinkedInWhatsAppE-Mail

    Related Articles

    A stylized five made of butterflies – visual for Claude Fable 5
    June 9, 20266 min

    Claude Fable 5 & Mythos 5: When AI Shifts from Tasks to Responsibilities

    Anthropic launches Claude Fable 5 and Mythos 5 – SOTA on almost all benchmarks. More interesting than the numbers: The s

    Read more
    OpenClaw Pricing Shock: How to Avoid the $500 Bill
    April 5, 20262 min

    OpenClaw Pricing Shock: How to Avoid the $500 Bill

    Anthropic just killed third-party tool coverage under Claude subscriptions. If you're running OpenClaw without prep, you

    Read more
    From Chat to Workflow: How Anthropic Is Turning Claude Into a Digital Coworker
    March 30, 20262 min

    From Chat to Workflow: How Anthropic Is Turning Claude Into a Digital Coworker

    Dispatch, Computer Use, persistent tasks – Anthropic is layering capabilities in an order that's no accident. A strategi

    Read more
    Why We Switched from ChatGPT to Claude – and What We Learned About LLMs Along the Way
    February 20, 20265 min

    Why We Switched from ChatGPT to Claude – and What We Learned About LLMs Along the Way

    We worked with ChatGPT for 18 months – then switched to Claude. Here's our honest comparison of all major LLMs and why C

    Read more
    Visualization of a large pale neural network sphere and a smaller bright sphere in cyan/yellow – the shrinking frontier of open models
    June 8, 20265 min

    Nex-N2-Pro: How the Open-Model Frontier Shrunk 75 % in Six Weeks

    Six weeks ago, DeepSeek-V4-Pro with 1.6 trillion parameters was the largest open-weight model ever released. Today, Nex-

    Read more
    Why 🦞 Became the Secret Handshake of the Agentic AI Movement
    May 19, 20263 min

    Why 🦞 Became the Secret Handshake of the Agentic AI Movement

    How a crustacean became the tribal emoji of the agentic AI scene – from Anthropic memes to X bios full of lobster claws.

    Read more
    Visualization of Kimi K2.6 long-horizon agents: a Moonshot crescent symbol alongside distributed sub-agent nodes over a coordination gridDeep Dive
    April 21, 20268 min

    Kimi K2.6: The Most Interesting AI Optimization in 2026 Isn't Intelligence – It's Duration

    Moonshot AI open-sourced Kimi K2.6 yesterday. 1 trillion parameters, 300 sub-agents, 13 hours of autonomous code refacto

    Read more
    Editorial illustration of the Claude Design launch – warm sand-tone background with the rust-orange Claude spark motif, glassmorphic UI panels showing a wireframe, color tokens, and a dashboard mockup, with subtle Adobe-red and Figma-purple accents hinting at the market disruption.
    April 17, 20265 min

    Claude Design Is Here: How Anthropic Labs Wiped $30B Off Figma, Adobe and Wix in a Single Day

    On April 17, 2026, Anthropic launched Claude Design – the first Anthropic Labs product for visual work. Powered by Opus

    Read more
    Claude Opus 4.7 Is Here: What Premium Teams Need to Know About the Tokenizer, xhigh, and Spend Controls
    April 17, 20265 min

    Claude Opus 4.7 Is Here: What Premium Teams Need to Know About the Tokenizer, xhigh, and Spend Controls

    Anthropic just released Claude Opus 4.7. Same price as 4.6, but noticeably better at coding, agents, and visual output.

    Read more