Safety and alignment

Anthropic introduces Constitutional AI training method

Anthropic introduced Constitutional AI, using written principles, model self-critique, revisions, and reinforcement learning from AI feedback to improve assistant harmlessness.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM52confidence 90/100

Why it moved the index

The method directly targeted scalable control of model behavior and was later confirmed by Anthropic as part of deployed Claude training. No exact December Claude checkpoint was publicly established, so no model profile is invented.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

AUDIT TRAIL

Assessment history

  1. R1
    Away 52 · confidence 90

    New December 2022 safety result with separate later primary evidence of practical deployment in Claude systems.

    12 Aug 2026