<aside> <img src="/icons/reorder_gray.svg" alt="/icons/reorder_gray.svg" width="40px" />

Table of Contents

</aside>


【0.】What changed in Opus 4.8

Anthropic shipped Claude Opus 4.8 on May 28, 2026. They call it a “modest but tangible” upgrade — and on raw benchmarks they’re right (GPT-5.5 still wins on terminal work). The real shift is quieter: the model now tells you when it isn’t sure. It’s about 4× less likely than 4.7 to let a flaw in its own code pass without flagging it.


【1.】The 4 Upgrades That Matter

<aside> 🔹

1. Honesty — it stops bluffing


<aside> 🔹

2. Effort control — you set how hard it thinks


<aside> 🔹

3. Dynamic workflows — many subagents, one run


<aside> 🔹

4. Fast mode — faster, and cheaper than it used to be



【2.】The Numbers

Capability Opus 4.7 Opus 4.8
Agentic coding 64.3% 69.2%
Multidisciplinary reasoning (with tools) 54.7% 57.9%
Agentic computer use 82.8% 83.4%
Computer use (Online-Mind2Web) 84%
Knowledge work (score) 1753 1890

Source: Anthropic, “Introducing Claude Opus 4.8,” May 28, 2026. Benchmarks are Anthropic’s own.

<aside> ▪️

Where it still falls short (the honest part)


  1. Incremental. Anthropic itself calls it “modest but tangible.” Happy on 4.7? No need to rush.
  2. GPT-5.5 still leads on terminal-heavy coding, advanced math, and web research.
  3. Prompt-injection resistance regressed slightly at the model level vs 4.7; product safeguards mostly compensate.
  4. Easy to overspend — a single max-effort + dynamic-workflow session can eat days of a small plan’s budget.
  5. Fast mode has no published accuracy benchmarks. Don’t use it for mission-critical work. </aside>

【3.】10 Mega Prompts for Opus 4.8