Cognition 🇺🇸 · SWE

SWE-2

Cognition's proprietary coding model for Devin reports near-frontier performance on FrontierCode 1.1 at substantially lower cost. It is deployed in Devin Desktop and CLI, with wider Web and Fusion availability planned through a staged rollout.

DOOM SCORE76.9out of 100model risk profile, not the overall index
0 comments · 0 votesOpen discussion

Public discussion is readable by everyone. Sign in to comment, reply, or vote.

No comments yet. Start the discussion.

CURRENT ASSESSMENT · REVISION 1

Why this model scores 76.9

High coding capability and integration into an autonomous software agent support strong capability and autonomy ratings. Deployment is meaningful but still rolling out. Misuse and control scores remain lower because the release provides provider-run trustworthiness tests but no independent dangerous-domain evaluation or observed loss of control.

Capability93
Autonomy88
Deployment79
Misuse potential63
Control difficulty50
MODEL-ATTRIBUTED EVIDENCE

News tied to SWE-2

The model score of 76.9 rates this model's risk profile. The overall Doom Index of 67.7 measures the complete temporally weighted evidence record. These values answer different questions.

NET MODEL-ATTRIBUTED INDEX CONTRIBUTION+0.19

Each article's current Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores use a bounded temporal offset around their technical profile, but never feed back into the overall index.

AUDIT TRAIL

Model score history

  1. R1
    Doom Score 76.9

    Creates the source-verified SWE-2 record from Cognition's dated launch, benchmark, cost, deployment, and trustworthiness details.

    11 Sept 2026
SHARE THE FINDINGS

Share this page

DoomBench social sharing card for SWE-2.
  1. SWE-2 by Cognition has a DoomBench model risk score of 76.9 out of 100, based on five transparent version-specific dimensions rather than the overall index.

  2. SWE-2's highest current DoomBench dimension is capability at 93.0 out of 100; the profile publishes every component score and its editorial rationale.

  3. DoomBench links 1 source-backed evidence item to SWE-2, while keeping the model's risk profile separate from each item's contribution to the live Doom Index.

    https://www.doombench.com/models/cognition-swe-2