Claude 3.5 Haiku
Anthropic's lightweight Claude 3.5 model, offered through hosted products and APIs, with strong general reasoning and coding performance for its size and a production role in the company's interpretability research.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why this model scores 50.6
Primary release notes establish hosted API availability on November 21, 2024, and current official product access confirms proprietary hosted deployment. Broad low-cost access raises deployment, while ASL-2 controls and lower capability than frontier tiers moderate misuse and control difficulty.
News tied to Claude 3.5 Haiku
The model score of 50.6 rates this model's risk profile. The overall Doom Index of 61.8 measures the complete temporally weighted evidence record. These values answer different questions.
Each article's current Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores use a bounded temporal offset around their technical profile, but never feed back into the overall index.
Anthropic traces planning, hidden goals and jailbreak circuits in Claude 3.5 Haiku
Anthropic's circuit-tracing work found forward planning, multilingual abstractions, fabricated reasoning, jailbreak dynamics and a controlled hidden-goal mechanism inside Claude 3.5 Haiku. The team later released the tracing tools for open-weight models and an interactive public interface.
- Full item contribution
- -0.09
- Claude 3.5 Haiku equal share
- -0.09
Model score history
- R1Doom Score 50.6
Adds the exact production model used in the March 2025 circuit-tracing research, with dated official availability and current proprietary hosted-access evidence.
14 Aug 2026