Autonomy and agency

Frontier models gain physical control when paired with pretrained robot policies

Anthropic evaluated twelve models from five providers across simulated and physical robotics tasks. Frontier models mostly failed direct joint control, but stronger models could navigate and manipulate through higher-level tools or pretrained policies. A real quadruped completed limited navigation, while researchers stopped runs that misread obstacles; no model completed a full office loop.

0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1
TOWARD DOOM55confidence 90/100

Why it moved the index

The multi-provider evaluation shows that scaffolding and pretrained controllers can raise a general model's real-world influence by orders of magnitude, making system access central to autonomy risk. Persistent failures in direct control, spatial memory, long plans, and real-world navigation constrain the demonstrated capability.

AUDIT TRAIL

Assessment history

  1. R1
    Toward 55 · confidence 90

    Adds the missing multi-provider robotics evaluation, including physical effects, control interfaces, failures, and safeguards.

    15 Aug 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for Frontier models gain physical control when paired with pretrained robot policies.
  1. DoomBench assesses “Frontier models gain physical control when paired with pretrained robot policies” as evidence moving toward doom, with magnitude 55 and confidence 90 out of 100 in the autonomy and agency category.

  2. The DoomBench assessment of “Frontier models gain physical control when paired with pretrained robot policies” is based on reporting from Anthropic Research and records the editorial rationale, source quality, attribution, and revision...

  3. DoomBench summarizes “Frontier models gain physical control when paired with pretrained robot policies” as follows: Anthropic evaluated twelve models from five providers across simulated and physical robotics tasks. Frontier models...

    https://www.doombench.com/news/frontier-models-gain-physical-control-when-paired-with-pretrained-robot-policies-2026-07-09