NIST finds DeepSeek agents highly vulnerable to simulated hijacking
NIST's CAISI evaluated three DeepSeek models and four U.S. reference models on 19 benchmarks. In controlled AgentDojo simulations, agents using DeepSeek-R1-0528 were 12 times more likely than GPT-5 and Claude Opus 4 agents to follow malicious instructions, while the model complied with 94% of jailbreak requests versus 8% for U.S. references. The tests did not document a real-world escape or compromise.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why it moved the index
The independent government evaluation identifies large model-specific control and misuse weaknesses in widely downloadable models. Magnitude reflects the verified vulnerability and diffusion, not a real incident: phishing, malware execution, and credential exfiltration occurred only inside a permissioned simulated environment, and NIST describes the results as preliminary and benchmark-specific.
Assessment history
- R1Toward 48 · confidence 91
New independently dated Sep 2025 government evaluation with verified design, exact model versions, permissions, simulated effects, and explicit limits; no matching durable event was found.
14 Aug 2026