Safety and alignment

Anthropic open-sources Petri for automated alignment audits

Anthropic released Petri to automate broad alignment audits; it was already used in Claude system cards and UK AISI testing, and later became part of every Claude alignment assessment from Sonnet 4.5 onward.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
AWAY FROM DOOM40confidence 90/100

Why it moved the index

An open, operational auditing framework already used by a frontier lab and a national evaluator directly expands practical capacity to detect deception, self-preservation, power-seeking, and reward hacking before deployment.

Assessment history

  1. R1
    Away 40 · confidence 90

    Adds previously missing primary-source evidence of a safety tool with separately verified practical adoption in frontier-lab and government evaluations.

    14 Aug 2026

Share this page

DoomBench social sharing card for Anthropic open-sources Petri for automated alignment audits.
  1. DoomBench assesses “Anthropic open-sources Petri for automated alignment audits” as evidence moving away from doom, with magnitude 40 and confidence 90 out of 100 in the safety and alignment category.

  2. The DoomBench assessment of “Anthropic open-sources Petri for automated alignment audits” is based on reporting from Anthropic and records the editorial rationale, source quality, attribution, and revision history.

  3. DoomBench summarizes “Anthropic open-sources Petri for automated alignment audits” as follows: Anthropic released Petri to automate broad alignment audits; it was already used in Claude system cards and UK AISI testing, and later became...

    https://www.doombench.com/news/anthropic-open-sources-petri-for-automated-alignment-audits-2025-10-06