Frontier models gain physical control when paired with pretrained robot policies
Anthropic evaluated twelve models from five providers across simulated and physical robotics tasks. Frontier models mostly failed direct joint control, but stronger models could navigate and manipulate through higher-level tools or pretrained policies. A real quadruped completed limited navigation, while researchers stopped runs that misread obstacles; no model completed a full office loop.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The multi-provider evaluation shows that scaffolding and pretrained controllers can raise a general model's real-world influence by orders of magnitude, making system access central to autonomy risk. Persistent failures in direct control, spatial memory, long plans, and real-world navigation constrain the demonstrated capability.
Assessment history
-
R1
Toward 55 · confidence 90
Adds the missing multi-provider robotics evaluation, including physical effects, control interfaces, failures, and safeguards.
15 Aug 2026