WideSWE: Can Coding Agents Coordinate Changes Across Repositories?
WideSWE tests coding agents on 120 cross-repository tasks; Codex CLI with GPT-5.6-sol succeeds on 42.50%.
WideSWE evaluates whether coding agents can coordinate changes across repositories, using 120 real tasks mined from 103 software ecosystems and balanced as 60 bug fixes and 60 features. Prompts come from related issues and pull requests, and hidden tests were adapted to accept diverse correct implementations while preserving required behavior and regression checks. Across seven agent configurations, full-task success ranges from 10.83% to 42.50%, with Codex CLI paired with GPT-5.6-sol highest. Agents miss needed changes, leave recognized edits unfinished, or modify the required repositories without satisfying the request; joint execution can use related-repo context, while independent per-repo runs mainly recover omitted work.
- 120 real tasks from 103 ecosystems, split evenly between bug fixes and features.
- Seven agent configurations fully succeed on 10.83% to 42.50% of tasks.
- Codex CLI with GPT-5.6-sol is the strongest configuration.
- Agents miss changes, leave edits unfinished, or fail to satisfy cross-repo requests.
- Joint execution uses related repositories; separate runs mainly recover omitted work.
Full article203 words · extracted from huggingface.co · click to collapse
Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multiple repositories. We introduce WideSWE to evaluate coding agents on such cross-repository tasks. Mining and reviewing changes across 103 software ecosystems yields 120 real-world tasks, balanced between 60 bug fixes and 60 features. We derive prompts from related issues and pull requests. We systematically review and adapt hidden tests to support diverse correct implementations while preserving required behavior and regression checks. Across seven agent configurations, full task success ranges from 10.83% to 42.50%, with the configuration pairing Codex CLI with GPT-5.6-sol achieving the highest rate. Trajectories show agents failing to identify necessary changes, recognizing changes but leaving them unfinished, or modifying the required repositories without fully satisfying the request. To examine whether working on one repository at a time can alleviate these difficulties, we compare it with joint execution under identical prompts. Independent execution mainly recovers omitted work and is less effective at correcting previously attempted but unsuccessful implementations. Joint execution can use information from related repositories to guide implementation and verification. Code is available at https://github.com/ZJU-ACES-ISE/WideSWE.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.33382