SWE-2
Cognition's proprietary coding model for Devin reports near-frontier performance on FrontierCode 1.1 at substantially lower cost. It is deployed in Devin Desktop and CLI, with wider Web and Fusion availability planned through a staged rollout.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
Why this model scores 76.9
High coding capability and integration into an autonomous software agent support strong capability and autonomy ratings. Deployment is meaningful but still rolling out. Misuse and control scores remain lower because the release provides provider-run trustworthiness tests but no independent dangerous-domain evaluation or observed loss of control.
News tied to SWE-2
The model score of 76.9 rates this model's risk profile. The overall Doom Index of 67.7 measures the complete temporally weighted evidence record. These values answer different questions.
Each article's current Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores use a bounded temporal offset around their technical profile, but never feed back into the overall index.
Cognition releases SWE-2 near frontier coding performance at sharply lower cost
Cognition released SWE-2 for Devin, reporting a 50 percent score on FrontierCode 1.1, within one point of Anthropic's Fable 5.1, at 64 percent lower cost. The company says the model was trained with reinforcement learning at multi-trillion-parameter scale and is available in Devin Desktop and CLI.
- Full item contribution
- +0.19
- SWE-2 equal share
- +0.19
Model score history
-
R1
Doom Score 76.9
Creates the source-verified SWE-2 record from Cognition's dated launch, benchmark, cost, deployment, and trustworthiness details.
11 Sept 2026