Anthropic πŸ‡ΊπŸ‡Έ Β· Claude Sonnet

Claude Sonnet 4.5

Claude Sonnet 4.5 is Anthropic's September 2025 closed, hosted frontier model for coding, agents, and computer use. Anthropic reported sustained focus beyond 30 hours, 77.2% on SWE-bench Verified, and 61.4% on OSWorld, alongside improved alignment and an ASL-3 deployment posture.

DOOM SCORE72.5out of 100model risk profile, not the overall index
0 comments Β· 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT Β· REVISION 1

Why this model scores 72.5

The release sits above Claude Sonnet 4 on coding, computer use, tool use, and sustained multi-step work, while broad API and product access makes deployment high. Its ASL-3 system card reports improved alignment and prompt-injection defenses, below-ASL-4 results, and no catastrophic cyber threshold, moderating misuse and control-difficulty assessments. The 30-hour focus claim supports elevated autonomy without implying fully independent open-ended operation.

Capability82
Autonomy77
Deployment92
Misuse potential60
Control difficulty48
MODEL-ATTRIBUTED EVIDENCE

News tied to Claude Sonnet 4.5

The model score of 72.5 rates this model's risk profile. The overall Doom Index of 61.1 measures the complete temporally weighted evidence record. These values answer different questions.

NET MODEL-ATTRIBUTED INDEX CONTRIBUTION+0.29

Each article's current Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores use a bounded temporal offset around their technical profile, but never feed back into the overall index.

Autonomy TOWARD

Anthropic releases Claude Sonnet 4.5 with longer autonomous task performance

Anthropic released Claude Sonnet 4.5 across its API and products, reporting state-of-the-art coding and computer-use results plus sustained focus for more than 30 hours on complex multi-step tasks. The accompanying system card places the model at the 2-to-8-hour software-engineering threshold while documenting ASL-3 safeguards and controlled alignment evaluations.

Full item contribution
+0.29
Claude Sonnet 4.5 equal share
+0.29
Read assessment β†’
AUDIT TRAIL

Model score history

  1. R1
    Doom Score 72.5

    Initial exact-version profile from Anthropic's dated launch and Claude Sonnet 4.5 system card.

    14 Aug 2026