Safety and alignment

Anthropic expands model-safety jailbreak bounties

Anthropic opened an invite-only bug bounty paying up to $15,000 for universal jailbreaks that could bypass forthcoming safeguards against high-risk chemical, biological, radiological, nuclear, and cybersecurity assistance.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM32confidence 92/100

Why it moved the index

Paying independent researchers for broad safeguard bypasses creates a practical discovery channel for severe model-control failures before wider deployment, though the program was initially limited in participation and scope.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

AUDIT TRAIL

Assessment history

  1. R1
    Away 32 · confidence 92

    New August 2024 advanced-model safeguard and vulnerability-research program.

    12 Aug 2026