Qwen2.5-VL-72B-Instruct
READER SUMMARYAn open 72-billion-parameter multimodal instruction model with long-video understanding, structured visual outputs and explicit visual-agent support for computer and phone use.
Why this model scores 76.5
Capability advances over Qwen2-VL in documents, video and general multimodal tasks. Direct computer and phone control materially raises autonomy. Open weights and multiple hosting channels drive near-maximal deployment, broaden misuse access and make downstream control difficult despite repository licenses and deployment-side safeguards.
0 comments · 0 votes
Sign in to join the discussion →
No comments yet. Start the discussion.
News tied to Qwen2.5-VL-72B-Instruct
The model score of 76.5 rates this model's risk profile. The overall Doom Index of 63.2 measures the recalibrated complete evidence corpus. These values answer different questions.
Each article's leave-one-out Doom Index contribution is divided equally among the exact models named on that article. This prevents multi-model evidence from being claimed in full on several model pages. Model risk scores are recalculated from their technical profile and related news, but never feed back into the overall index.
Qwen releases open visual agents for computer and phone use
Alibaba's Qwen team released Qwen2.5-VL models at 3B, 7B and 72B scales, with open base and instruction weights and direct visual-agent support for computer and phone use.
- Full item contribution
- +0.04
- Qwen2.5-VL-72B-Instruct equal share
- +0.04
Model score history
- R1Doom Score 76.5
Initial exact source-backed profile for the flagship Qwen2.5-VL 72B instruction tier.
12 Aug 2026