OSWorld-Pro: Process-based Evaluation for Computer Use Agents
OSWorld-Pro scores computer-use agents on over 300 tasks; Claude Opus 5 reaches 75.7%.
OSWorld-Pro is a process-based benchmark for computer-use agents with more than 300 tasks, over 2,800 subgoals, and more than 67,000 human annotations. Unlike OSWorld’s end-state functional checks, human-aligned LLM judges score fulfillment of sequentially dependent subgoals. Even strong models struggle: Claude Opus 5 reaches 75.7% on OSWorld-Pro versus 83.4% on OSWorld. The authors highlight process failures such as subgoal-irrelevant actions and click-based GUI mistakes.
- More than 300 tasks and 2,800 subgoals, backed by over 67,000 human annotations.
- Human-aligned LLM judges score sequentially dependent subgoals, not only final deliverables.
- Claude Opus 5 scores 75.7% on OSWorld-Pro versus 83.4% on OSWorld.
- Failure modes include subgoal-irrelevant actions and click-based GUI mistakes.
- Keyboard-input errors and click errors imply different mitigation strategies.
Full article184 words · extracted from huggingface.co · click to collapse
Evaluation of Computer-Use Agents (CUAs) is often limited to the final deliverables they create (at the end of hundreds of steps) and assessed with functional verifiers, as seen in OSWorld. However, such evaluation of end-state performance lacks transparency into how and why agents fail in various tasks, obfuscating critical insight for subsequent improvement. For instance, agents that err during keyboard inputs would require a different mitigation strategy from those that fail to precisely provide click-based inputs on the graphical UI. We introduce OSWorld-Pro: a set of over 300 tasks containing over 2800 subgoals to enable the procedural evaluation of CUAs grounded in over 67,000 human annotations. We use robust human-aligned LLM-Judges to evaluate the fulfillment of OSWorld-Pro subgoals and thereby reveal the progress that models make throughout a series of sequentially dependent subgoals. Our findings reveal that OSWorld-Pro is challenging even for state-of-the-art LLMs, with top performers like Claude Opus 5 achieving only 75.7% vs. 83.4% on OSWorld. Furthermore, we identify critical process-focused failure modes of various models (e.g. subgoal-irrelevant actions and click-based mistakes) to provide insights to improve performance and efficiency of CUAs.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.24890