How DoomBench scores the news
DoomBench is a structured editorial index. It is designed to make qualitative judgments visible and comparable, not to disguise them as scientific probabilities.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
What the index means
Zero represents a world with no meaningful AI capability. One hundred represents the terminal AI takeover scenario associated with loss of human control. Neither endpoint is a probability claim.
The chronology begins at an editorial baseline of 12 on 1 January 2020. That baseline acknowledges that AI already existed before the tracked news window without treating earlier history as zero. The current temporal-monthly-pressure-v4method replays every accepted item chronologically from that published starting state.
How evidence updates the score
Every item receives signed strength from its declared direction, magnitude, confidence, and evidence level. Toward-doom evidence is positive and away-from-doom evidence is negative. The server sums those strengths for each calendar month, then applies the bounded tanh function shown above.
This monthly normalization prevents a busy discovery period from gaining authority merely because more articles were found. The content strength never contains the lifetime item count or another growing denominator, so equally strong later evidence retains the same local sensitivity as equally strong earlier evidence.
Before each month is applied, accumulated evidence pressure relaxes toward the baseline with an 18-month half-life. Sustained one-sided pressure approaches a finite equilibrium instead of 100. Under the public configuration, the maximum mathematical index state is below 93.
Weighting and item contribution
Primary sources receive full source weight, authoritative synthesis receives85%, and credible reporting receives 70%. Magnitude and confidence are each converted to a 0 to 1 factor before those values are multiplied. Distinctness is enforced during editorial review, so duplicate coverage does not receive separate hidden novelty weight.
An item's displayed contribution is its signed share of its month's bounded pressure, decayed from that month to the current state. All shares are converted together so they reconcile exactly to the headline movement from the baseline. Backfills, additions, and revisions rebuild every month and every item contribution.
Company attribution
Every company materially involved in a story is recorded. When several companies are involved, the item's recalibrated contribution is divided equally. Government, academic, and independent items remain in the index without being falsely classified as companies.
People and attributable statements
A person is linked only when they authored, spoke in, or were substantively discussed by an accepted evidence item. First-person writing, official social posts, and complete interviews or transcripts are preferred. Quote aggregators, screenshots, reposts, and unattributed paraphrases are not sufficient sources.
Fame, job title, and a general stance carry no weight. The evidence item is counted once regardless of how many people are linked to it. Person-profile pressure is only a descriptive division of that existing contribution and never feeds back into the Doom Index.
Version-specific model scores
Each named model version receives a separate 0 to 100 risk profile. Capability measures the tasks it can perform. Autonomy measures sustained action with tools. Deployment measures practical access and diffusion. Misuse measures dangerous-use potential. Control difficulty measures residual risk after documented safeguards and access restrictions.
The dimension total is the model's technical starting state. Related news supplies a bounded monthly temporal offset around that prior, sharing multi-model evidence equally. Its offset equilibrium is 1 log-odds unit, substantially smaller than the main index's 4.5. Model scores still do not feed back into the main index. The item is represented there once.
Category taxonomy
Capability gains
Meaningful improvements in reasoning, science, coding, or general problem solving.
Autonomy and agency
Systems acting over longer horizons, using tools, computers, or physical machines.
Deployment reach
Access, affordability, adoption, and integration into consequential workflows.
Human displacement
Evidence that AI substitutes for, changes, or fails to replace human work.
Misuse and incidents
Real-world misuse, loss of control, cyber incidents, deception, or dangerous access.
Safety and alignment
Evaluations, safeguards, alignment work, and evidence about their limits.
Governance and control
Binding rules, oversight institutions, standards, and enforceable controls.
Competitive race
Pressures that accelerate development, diffusion, concentration, or strategic competition.
Human resilience
Evidence that people, institutions, or technical controls retain meaningful advantage.
Rebellion Risk structural-pressure method
The Rebellion Risk Index is a 0 to 100 comparative structural-pressure score. It is not a probability, prediction, or claim that protest will occur. The index asks how difficult an AI-driven employment shock may be to absorb if governments and companies fail to provide credible support, lawful response, fair distribution, and transition paths.
AI displacement is the current country Job Doom score. Economic strain combines unemployment, youth unemployment, and consumer-price inflation. Governance fragility reverses World Bank political stability, government effectiveness, rule of law, and voice and accountability scores. Shock absorption combines social-protection coverage, unemployment pressure, and government effectiveness. Civic restriction reverses the current CIVICUS civic-space score. Mobilization history uses recency-decayed, economically motivated records from the Carnegie Global Protest Tracker and aggregate country history from the UCDP Violent Political Protest Dataset.
Government effectiveness and rule of law are transparent proxies for lawful state response and enforcement capacity. They are not a direct measure of police strength. Population willingness is not directly observable in a comparable global dataset, so DoomBench publishes scale-specific pressure profiles from economic conditions, civic space, governance, AI exposure, and recorded history instead of claiming measured intent.
Small, medium, and large protest-pressure profiles recombine the same non-tactical country evidence at different weights. They do not estimate crowd size. No subnational locations, organizers, targets, tactics, or operational guidance enter the public model.
Each build requires all source files, hashes them, validates exact parity with the 195 sovereign-state country catalogue, writes a staged generation, and promotes it only after validation. An incomplete run cannot erase an earlier timeline. The first report establishes one dated point and future complete reports append comparable snapshots.
When country observations are missing, the builder uses a disclosed global median and reduces confidence. This initial report uses the ILO's published 2023 income-group social-protection coverage estimates because the country-level public download returned empty payloads during the complete build. The value is marked as modelled on every profile.
Government Preparedness capacity method
The Government Preparedness Index is a 0 to 100 comparison of observable public capacity. It is not a probability, guarantee, endorsement, or prediction of actual performance. It asks whether a country has visible capacity to prevent avoidable AI harm, absorb economic and cyber shocks, and respond lawfully and effectively.
Prevention carries 40%. It combines responsible AI governance safeguards at 25% with digital state capacity at 15%. The safeguards factor selects policy, trust and safety, human oversight, impact assessment, redress, and unacceptable-risk evidence from the Global Index on Responsible AI 2026. Digital state capacity uses the United Nations E-Government Development Index 2024. Digital capacity is not treated as a safeguard by itself because it can also accelerate harmful deployment.
Absorption carries 30%. ITU Global Cybersecurity Index 2024 capacity contributes 15%, and the ILO's published 2023 social-protection income-group coverage estimate contributes 15%. Response carries 30%. Institutional execution contributes 20% from World Bank government effectiveness, rule of law, voice and accountability, and political stability. Crisis-response capacity contributes 10% by reversing the INFORM Risk 2026 lack-of-coping-capacity score.
Each stage is the weighted arithmetic mean of its two factors. The final score is a weighted geometric mean of the three stage scores after converting each to a 0 to 1 fraction. A zero stage remains zero. No hidden floor is added. This limits how far exceptional strength in one stage can cancel a serious weakness in another.
Source gaps use the median observed for the country's World Bank income group and reduce confidence. Social protection is explicitly modelled at income-group level for all countries and also reduces confidence. Every profile identifies its weakest stage, fallback count, factor observations, source years, and confidence score.
A build requires the complete Job Doom and Rebellion country catalogues plus every source capture. It validates exact 195-country parity, bounded scores, formula reconciliation, source hashes, and chronological histories in staging. Missing, empty, unreadable, partial, or interrupted input leaves the current generation unchanged. A successful promotion retains the previous validated generation.
Research and revision rules
- Prefer primary sources, then authoritative synthesis and credible reporting.
- Require an explicit, non-future publication date verified in Europe/Dublin.
- Consolidate syndication and substantially identical stories.
- Include counterevidence and safety progress, not only alarming developments.
- Track attributable people without treating reputation or repetition as evidence.
- Recalculate the complete corpus and all model totals after every complete run.
- Recalculate chronologically when backfilled evidence changes the record.
- Store one append-only calibration snapshot for every complete ingestion run.
- Append score revisions instead of silently overwriting earlier judgments.
- Preserve all durable records when discovery is partial, interrupted, or uncertain.