SolveEdit: Benchmarking Visual Problem Solving in Generative Models
SolveEdit benchmarks goal-driven visual edits; SolveEdit-PLAN lifts GPT-Image-2 from 57.0% to 71.6%.
SolveEdit benchmarks goal-driven visual problem solving: given an image and a goal, a model must infer a valid scene transformation and execute it while preserving unrelated content. The set contains 2,728 cases, and atomic transition contracts let SolveScore measure completion and unintended changes without a single reference output. The strongest evaluated model scores only 57.0% SolveScore. SolveEdit-PLAN, a two-stage visual planner that instantiates the transition before generation, improves SolveScore by 9.1 points on average across three generators, including a rise from 57.0% to 71.6% for GPT-Image-2, without modifying the editor.