Safety and alignment

Anthropic deploys safeguards against autonomous election influence operations

Anthropic reported always-on classifiers, monitoring, system prompts, and election controls that caused safeguarded models to refuse nearly all autonomous influence-operation tasks despite strong raw capability.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM48confidence 78/100

Why it moved the index

Deployed controls substantially reduced autonomous political-manipulation behavior in model-specific tests. Confidence is limited because the results are developer-reported and not an independent real-election outcome study.

AUDIT TRAIL

Assessment history

  1. R1
    Away 48 · confidence 78

    New deployed safeguard and evaluation evidence directly addressing autonomous manipulation risk.

    11 Aug 2026