Safety and alignment

Anthropic deploys stronger browser-agent prompt-injection defenses

Anthropic combined Claude Opus 4.5 training, content classifiers, intervention logic, and continuous red teaming to reduce adaptive prompt-injection attack success to about one percent and expanded Claude for Chrome to beta.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM35confidence 85/100

Why it moved the index

A measured and deployed defense against agent hijacking directly strengthens human control over browser agents that process attacker-controlled content.

AUDIT TRAIL

Assessment history

  1. R1
    Away 35 · confidence 85

    New historical safeguard-deployment record supported by Anthropic's dated technical report.

    11 Aug 2026