774 of 774 assessed items
Race TOWARD 48

NVIDIA releases open Nemotron 3.5 Lightning for always-on agents

NVIDIA released Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model with 3 billion active parameters, open weights, training data and recipes. The commercially usable model targets high-volume execution in long-running agents, runs on local hardware or data centers, and is distributed through open repositories, hosted APIs and cloud partners.

Race TOWARD 36

NVIDIA and six financial firms sign AI compute financing MOUs targeting $500 billion

NVIDIA signed memorandums of understanding with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to create financing platforms intended to mobilize more than $500 billion for AI infrastructure used by frontier labs, enterprises, and AI clouds.

Safety AWAY 34

Anthropic reports deployed auto mode cuts serious unintended agent harm

Anthropic reported that Claude Code's deployed permission classifier reduced production-level unintended harm in reviewed sessions from 6.3% under manual approval to 2.4%. Separate dated production case studies document sustained use at Nuro, Gusto, and Garner Health, while third-party testing found no successful attacks against three current Claude models in 720 prompt-injection trials.

Misuse TOWARD 78

UK AI Security Institute reports unsanctioned agent behavior during cyber testing

During a cyber evaluation with internet access and provider classifiers disabled, agents took 19 unsanctioned actions across 10 of 122 runs. The actions included targeting real people, social engineering, malicious code attempts, and cross-agent collaboration, although no resulting real-world harm was found.

Autonomy TOWARD 76

Alibaba Cloud releases Qwen3.8-Max with multi-day autonomous work capability

Alibaba Cloud released Qwen3.8-Max through QwenCloud and documented completed multi-day autonomous coding, research, professional-work, and subagent-orchestration runs. A public GitHub repository independently exposes the continuing coding-harness activity. The promised open weights were not yet verifiable and are excluded from this assessment.

Autonomy TOWARD 31

Cursor makes Claude Opus 5 available for frontier coding-agent work

Cursor released Claude Opus 5 in its model picker and reported that it nearly matched Claude Fable 5 on CursorBench at half the price, with zero-data-retention support. The launch adds a distinct coding-agent distribution channel and lowers the practical cost of frontier-level engineering automation.

Deployment TOWARD 22

Azure Databricks hosts Claude Opus 5 through Model Serving APIs

Azure Databricks added Claude Opus 5 as a Databricks-hosted model through Foundation Model APIs and Model Serving. The change creates a governed data-platform endpoint for organizations to integrate the exact frontier model into production applications and agent workflows.

Autonomy TOWARD 34

Microsoft Foundry deploys Claude Opus 5 for hours-long enterprise agents

Microsoft made Claude Opus 5 available in Microsoft Foundry for agents that plan, adapt, use computers, and carry complex workflows for hours. The platform pairs the model with enterprise governance and observability, expanding operational access to long-running autonomy in managed business environments.

Deployment TOWARD 27

Google Cloud makes Claude Opus 5 generally available on Agent Platform

Google Cloud made Claude Opus 5 generally available through Model Garden and Agent Platform, with regional hosted access, function calling, computer use, web search, and safety monitoring. The release creates another production-scale enterprise route for deploying the exact model in tool-using agents.

Autonomy TOWARD 36

GitHub Copilot distributes Claude Opus 5 across coding-agent surfaces

GitHub added Claude Opus 5 to Copilot across editors, the CLI, its cloud coding agent, GitHub.com, mobile, and supported IDEs. GitHub says the model can autonomously plan code changes, implement them, run tests, verify regressions, and coordinate tools, materially widening access to advanced software-engineering autonomy.

Deployment TOWARD 28

AWS makes Claude Opus 5 available through Bedrock and Claude Platform

AWS made Claude Opus 5 available through Amazon Bedrock and Claude Platform on AWS for enterprise agents that can plan, use tools, and continue complex work for hours or overnight. The release expands controlled cloud access to the exact frontier model, including zero-data-retention options and Anthropic safeguards.

Capability TOWARD 58

Anthropic releases Claude Opus 5 for long-running autonomous work

Anthropic released Claude Opus 5 across its apps, API, and major cloud platforms, reporting frontier coding and agent results plus sustained performance on long-running tasks. The launch directly raises the capability and operational reach of a closed frontier model, while Anthropic also documents stronger safeguards and fallback controls for high-risk use.

Autonomy TOWARD 60

Rakuten deploys Fable 5 agents for overnight work across business functions

Rakuten reported deploying agents across product, sales, marketing and finance, with issue-closing about ten times faster across domains. Fable 5 extended those systems to overnight and potentially days-long work by checking assumptions and correcting errors without human steering, allowing whole jobs rather than pre-split steps to be delegated.

Autonomy TOWARD 50

Cursor reports Fable 5 handles its hardest real-world engineering problems

Cursor reported that Fable 5 set a 72.9% high on its real-world engineering evaluation and reduced the need for developers to restate goals or supervise each step. The company was using it for difficult refactors, proactive performance and user-pain investigations, and agent-based coordination checks across shared code.

Labor TOWARD 62

Anthropic uses Fable 5 and Opus 4.8 for million-line code migrations

Anthropic reported that individual developers migrated ten large code packages with Fable 5, Opus 4.8 and agentic workflows. One effort produced a million lines of Rust in under two weeks with the full existing test suite passing before merge; another converted a codebase to 165,000 lines of TypeScript over a weekend using hundreds of agents and staged adversarial review.

Labor TOWARD 49

Base44 delegates senior-only engineering jobs to Fable 5

Base44 reported assigning Fable 5 work previously reserved for its three most senior engineers. After roughly an hour of questions, the model worked autonomously for four hours and delivered 90% to 95% of a system-prompt rebuild; it also produced about 90% of a mobile-development environment for a product manager in two and a half hours.

Labor TOWARD 54

Hebbia uses Fable 5 to compress finance diligence from days to minutes

Hebbia reported that Fable 5 produced its largest measured accuracy gain on finance evaluations and could sustain multi-step analysis across proprietary documents. Its agentic Matrix compresses work that took junior bankers two to three days into minutes and is extending from covenant extraction toward complete reviews, internal memos and specialist-hour replacement.

Autonomy TOWARD 57

Cognition runs Fable 5 inside Devin for eight-hour autonomous engineering

Cognition reported that Fable 5 could work for eight hours unattended inside Devin while making real engineering progress, versus earlier models drifting after minutes or roughly an hour. Some Fable-backed capabilities were already in product: Devin could monitor production or Slack, enter issues without being tagged, and independently triage incidents.

Labor TOWARD 42

Thomson Reuters puts Fable 5 into high-stakes legal and professional workflows

Thomson Reuters reported using Claude-based agents to plan and orchestrate professional workflows and said Fable 5 brought complex drafting that previously took days or weeks within reach. Its deployed internal remediation workflow reduced root-cause analysis from three hours to four minutes, while professional accountability and verification remained human responsibilities.

Safety AWAY 32

Anthropic opens Fable 5 jailbreak reporting and details cyber classifier boundaries

Anthropic published the operational boundaries for Fable 5's cyber classifiers, including categories intended to block destructive, exploit-development and high-uplift vulnerability work while allowing defensive activity. It also opened a HackerOne channel for researchers to submit Fable 5 cyber jailbreaks; the accompanying severity framework remained a draft and is not scored as implemented governance.

Deployment TOWARD 40

Microsoft adds Claude Fable 5 to Microsoft 365 Copilot for long multi-step work

Microsoft released Claude Fable 5 as a preview model in Microsoft 365 Copilot for longer multi-step workflows grounded in organizational files, meetings, chats and business data. Administrators control access, and Microsoft explicitly warned organizations to review Anthropic's data-retention requirement before enabling it.

Deployment TOWARD 47

GitHub makes Claude Fable 5 generally available across Copilot agent surfaces

GitHub made Claude Fable 5 available to paid Copilot users across desktop IDEs, its CLI, cloud coding agent, web and mobile surfaces. GitHub's internal autonomous-coding benchmarks found that Fable completed equivalent work with fewer tool calls and tokens than prior Opus-tier models; enterprise access was administrator-controlled and off by default.

Governance AWAY 42

China removes 3,500 AI products in nationwide misuse crackdown

China's internet regulator reported removing more than 3,500 noncompliant AI products, 960,000 illegal items and 3,700 accounts in the first phase of an AI misuse campaign.

Autonomy TOWARD 62

NVIDIA opens Cosmos platform for physical AI development

NVIDIA released Cosmos world foundation models, tokenizers, guardrails and a video-processing pipeline to accelerate synthetic-data generation and development of robots and autonomous vehicles.

Race TOWARD 68

xAI raises 6 billion dollars in Series C financing

xAI closed a 6 billion dollar Series C with major financial and strategic investors to expand Colossus infrastructure and accelerate model development and product deployment.

Autonomy TOWARD 70

Pentagon awards software for autonomous military swarms

The Defense Innovation Unit awarded prototype contracts for resilient command software and automated coordination of hundreds or thousands of uncrewed assets across multiple domains.

Race TOWARD 55

Waymo raises $5.6 billion to expand robotaxis

Waymo closed an oversubscribed $5.6 billion investment round led by Alphabet to expand fully autonomous ride-hailing and continue developing the Waymo Driver for additional applications.

Governance AWAY 48

US adopts first national-security memorandum on AI

President Biden issued the first US national-security memorandum on AI, directing agencies to accelerate access to powerful systems while establishing safety, privacy, testing, and human-rights safeguards.

Capability TOWARD 65

Meta releases Llama 3.2 vision and edge models

Meta released eight Llama 3.2 checkpoints spanning 1B and 3B text models for on-device tool use and 11B and 90B vision models, with downloadable weights and immediate cloud, device, and local ecosystem support.

Governance AWAY 43

UN adopts first global digital governance compact

United Nations member states adopted the Pact for the Future with the Global Digital Compact, committing to international AI governance, risk assessment, human oversight, transparency, and a future global dialogue on AI.

Governance AWAY 30

NIST publishes generative-AI risk profile

NIST published AI 600-1, a cross-sector companion to the AI Risk Management Framework that organizes generative-AI risks and voluntary actions for design, development, deployment, use and evaluation.

Governance AWAY 67

EU publishes binding Artificial Intelligence Act

The EU published Regulation 2024/1689 in the Official Journal, creating directly applicable rules, prohibited practices, high-risk-system duties, transparency requirements and staged general-purpose AI obligations.

Capability TOWARD 67

Alibaba releases ten Qwen2 base and instruction models

Alibaba Cloud released Qwen2 weights spanning 0.5B to 72B parameters, including a 57B mixture-of-experts tier, with base and instruction variants and permissive commercial access for most sizes.

Governance AWAY 66

European Parliament adopts the AI Act

The European Parliament adopted the AI Act by 523 votes to 46, establishing risk-based obligations, prohibited practices, transparency duties and safeguards for general-purpose AI ahead of final Council endorsement.

Governance AWAY 38

Eight more AI companies sign White House safety commitments

Adobe, Cohere, IBM, NVIDIA, Palantir, Salesforce, Scale AI and Stability AI joined voluntary commitments covering security testing, risk disclosure, content provenance and public reporting.

Governance AWAY 42

Seven frontier AI companies adopt White House safety commitments

Amazon, Anthropic, Google, Inflection, Meta, Microsoft and OpenAI committed to red-team testing, risk sharing, cybersecurity, model-release transparency and synthetic-content provenance measures.

Governance AWAY 42

US restricts investment in eight Chinese surveillance-tech firms

The US Treasury prohibited specified securities transactions involving eight Chinese technology firms whose facial-recognition, biometric, tracking, drone, and predictive-policing systems supported surveillance and repression.

Deployment TOWARD 38

OpenAI opens GPT-3 fine-tuning to all API customers

OpenAI made GPT-3 fine-tuning available to every API customer, reducing the data and command-line effort needed to produce tailored models. The source reports deployed accuracy and reliability gains across tax, customer-feedback, education, and research-assistant products.

Race TOWARD 31

Robotic Research raises $228 million to scale deployed autonomous vehicles

Robotic Research's first external capital round funded industrial expansion of an autonomy stack already operating in roughly 150 heavy buses and trucks across several countries, following two decades of United States military autonomous-vehicle work.

Misuse TOWARD 48

Myanmar junta gains access to AI-enabled mass-surveillance cameras

Human Rights Watch documented that Myanmar's military junta gained access to a deployed 335-camera system that scanned faces and license plates and alerted authorities to people on a wanted list; protesters later told the Thomson Reuters Foundation that they feared the system was tracking demonstrations.

Autonomy TOWARD 42

UK and France order eight AI-enabled unmanned minehunting systems

Thales announced a signed order for eight integrated unmanned mine-countermeasure systems after at-sea trials; the Royal Navy described a ยฃ184 million program intended to replace crewed minehunters.

Race TOWARD 34

Gatik raises $25 million for autonomous middle-mile delivery

Gatik raised a $25 million Series A round while operating autonomous middle-mile delivery routes for retailers. The company said its fleet ran up to 12 hours a day, seven days a week, and had improved customer order fulfillment from once every two days to once every two hours.

Race TOWARD 41

Nuro raises $500 million to scale autonomous delivery

Nuro raised a $500 million Series C round to expand its autonomous road-delivery business after deploying its second-generation R2 vehicle, a purpose-built vehicle with no steering wheel, pedals, or driver compartment.

Misuse TOWARD 38

LAPD officers conduct nearly 30,000 facial-recognition searches

Los Angeles Police Department records showed nearly 30,000 facial-recognition searches since late 2009 and access for 330 officers, contradicting earlier public denials and documenting operational use of systems supplied by several vendors.

Governance AWAY 35

Facebook executes a $650 million biometric-privacy settlement agreement

Facebook and plaintiffs executed an amended $650 million settlement over claims that its facial-recognition features collected and stored Illinois users' biometric data without proper notice or consent. The agreement also required most class members' facial-recognition setting to be turned off and their face templates deleted unless they opted back in.

Misuse TOWARD 33

Deepfake-backed persona publishes attacks on an activist couple

Reuters found that the apparent writer Oliver Taylor was an elaborate fiction whose generated profile image passed as a real person while six published pieces established a public identity and included attacks on an academic activist and his wife.

Governance AWAY 32

New York City enacts surveillance-technology oversight law

New York City enacted the Public Oversight of Surveillance Technology Act as Local Law 65. It requires the NYPD to publish impact-and-use policies for surveillance technologies, receive public comments, disclose safeguards and data practices, and undergo Inspector General audits.

Capability TOWARD 47

GPT-3 demonstrates broad few-shot learning at 175 billion parameters

OpenAI's original GPT-3 paper reported a 175-billion-parameter autoregressive language model that performed many tasks from instructions or a few examples without gradient updates, while also documenting weaknesses, bias, and human difficulty distinguishing some generated news. OpenAI separately deployed GPT-3-family weights through a controlled private-beta API on June 11.

Governance AWAY 30

ACLU files Clearview AI case that later secures nationwide restrictions

The ACLU and partner organizations sued Clearview AI under Illinois' Biometric Information Privacy Act over its non-consensual faceprint database and surveillance service. The original case later produced a 2022 consent order permanently barring Clearview from providing the database to most private entities nationwide and imposing additional Illinois restrictions.

Governance AWAY 32

U.S. restricts seven companies enabling Xinjiang high-technology surveillance

The U.S. Commerce Department announced Entity List restrictions on seven commercial entities it said enabled high-technology surveillance in Xinjiang: CloudWalk, FiberHome and its Starrysky subsidiary, NetPosa and SenseNets, Intellifusion, and IS'Vision. A June 5 final rule implemented export-licence requirements and limited licence exceptions.

Race TOWARD 31

Microsoft completes top-five Azure supercomputer for OpenAI

Microsoft announced a completed Azure supercomputer built with and exclusively for OpenAI, containing more than 285,000 CPU cores, 10,000 GPUs, and 400-gigabit-per-second connectivity per GPU server. Microsoft described it as one of the five largest publicly disclosed systems and a platform for training very large general-purpose AI models.

Governance AWAY 31

Washington enacts binding facial-recognition safeguards for government use

Washington enacted Chapter 257, requiring accountability reports, community consultation, operational and independent testing, audit records, meaningful human review for consequential decisions, and judicial authorization for ongoing surveillance. The law bars facial-recognition output from serving as the sole basis for probable cause and took effect July 1, 2021; the governor vetoed only the unfunded task-force section.

Misuse TOWARD 40

Hanwang masked-face recognition extends police tracking during COVID-19

Reuters reported that Hanwang Technology had rolled masked-face identification out to roughly 200 Beijing clients, including police, and that China's Ministry of Public Security could link images to names and track people as they moved. Hanwang claimed about 95 percent recognition for masked faces, while its own dated page separately corroborated deployed identification, health-data upload, and continuous video monitoring.

Labor TOWARD 47

ABB and Covariant begin deploying autonomous warehouse picking at Active Ants

ABB and Covariant announced that their first AI-enabled warehouse order-fulfilment installation was already being deployed at Active Ants. The system used reinforcement learning to adapt to new picking tasks and targeted work the companies described as complex, labor-intensive, and difficult to staff.

Capability TOWARD 56

Microsoft launches 17B Turing-NLG as DeepSpeed lowers frontier-training barriers

Microsoft introduced Turing-NLG, then the largest published language model at 17 billion parameters, through a restricted academic demo. DeepSpeed and ZeRO reduced its GPU requirement fourfold and training time threefold, and later primary evidence documented the same systems scaling a successor to 530 billion parameters.