Anthropic 🇺🇸 · Claude Opus

Claude Opus 4.8

A highly capable agentic model with stronger coding, computer use, long-running work, and documented improvements in alignment behavior.

DOOM SCORE83.9out of 100model risk profile, not the overall index
0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 6

Why this model scores 83.9

Near-frontier autonomy and dangerous-domain capability create substantial risk, partly offset by lower measured misalignment and deployed cyber safeguards.

Capability93
Autonomy92
Deployment75
Misuse potential84
Control difficulty68
MODEL-ATTRIBUTED EVIDENCE

News tied to Claude Opus 4.8

The model score of 83.9 rates this model's risk profile. The overall Doom Index of 67.9 measures the complete temporally weighted evidence record. These values answer different questions.

NET MODEL-ATTRIBUTED INDEX CONTRIBUTION+0.52

Each article's current Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores use a bounded temporal offset around their technical profile, but never feed back into the overall index.

Safety AWAY

Anthropic shows automated researchers can mitigate ten alignment failures

In a controlled study, Claude autonomously developed post-training methods that improved all ten tested alignment-failure categories without measured capability loss, generalized to withheld evaluations and larger models, and closed 65% of an early Claude Opus 4.8 checkpoint's measured safety gap in 60 hours.

Full item contribution
-0.11
Claude Opus 4.8 equal share
-0.06
Read assessment →
Autonomy TOWARD

Anthropic opens Model Hardware Standard preview for AI-controlled laboratory equipment

Anthropic opened a limited research preview of a model-agnostic hardware interface after controlled laboratory proofs showed AI agents coordinating physical instruments, adjusting parameters, recovering from some errors, and running closed-loop experiments. One Carnegie Mellon demonstration used Claude Opus 4.8 to control incompatible lab equipment and complete a corrected dose-response run without human input.

Full item contribution
+0.12
Claude Opus 4.8 equal share
+0.12
Read assessment →
Labor TOWARD

Anthropic uses Fable 5 and Opus 4.8 for million-line code migrations

Anthropic reported that individual developers migrated ten large code packages with Fable 5, Opus 4.8 and agentic workflows. One effort produced a million lines of Rust in under two weeks with the full existing test suite passing before merge; another converted a codebase to 165,000 lines of TypeScript over a weekend using hundreds of agents and staged adversarial review.

Full item contribution
+0.16
Claude Opus 4.8 equal share
+0.08
Read assessment →
AUDIT TRAIL

Model score history

  1. R6
    Doom Score 83.9

    Exact-version evidence chronology replayed after run doombench-20260829-065948-current-and-people-backfill under temporal-monthly-pressure-v4.

    29 Aug 2026
  2. R5
    Doom Score 84.0

    Exact-version evidence chronology replayed after run doombench-hourly-news-20260828-201012 under temporal-monthly-pressure-v4.

    28 Aug 2026
  3. R4
    Doom Score 83.9

    Source-backed model availability audit using the model's existing primary or authoritative catalogue evidence. Exact-version evidence chronology replayed under temporal-monthly-pressure-v4.

    12 Aug 2026
  4. R3
    Doom Score 83.9

    Exact-version evidence chronology replayed after run doombench-hourly-news-20260812-132112 under fixed-sensitivity-v3.

    12 Aug 2026
  5. R2
    Doom Score 84.1

    Full-corpus evidence recalculated after run intensive-backfill-20260811-191225 under bounded-corpus-v2.

    11 Aug 2026
  6. R1
    Doom Score 83.8

    Initial source-backed model assessment

    11 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for Claude Opus 4.8.
  1. Claude Opus 4.8 by Anthropic has a DoomBench model risk score of 83.9 out of 100, based on five transparent version-specific dimensions rather than the overall index.

  2. Claude Opus 4.8's highest current DoomBench dimension is capability at 93.0 out of 100; the profile publishes every component score and its editorial rationale.

  3. DoomBench links 4 source-backed evidence items to Claude Opus 4.8, while keeping the model's risk profile separate from each item's contribution to the live Doom Index.

    https://www.doombench.com/models/anthropic-claude-opus-4-8