SolveEdit: Benchmarking Visual Problem Solving in Generative Models
SolveEdit benchmarks goal-driven visual edits; SolveEdit-PLAN lifts GPT-Image-2 from 57.0% to 71.6%.
SolveEdit benchmarks goal-driven visual problem solving: given an image and a goal, a model must infer a valid scene transformation and execute it while preserving unrelated content. The set contains 2,728 cases, and atomic transition contracts let SolveScore measure completion and unintended changes without a single reference output. The strongest evaluated model scores only 57.0% SolveScore. SolveEdit-PLAN, a two-stage visual planner that instantiates the transition before generation, improves SolveScore by 9.1 points on average across three generators, including a rise from 57.0% to 71.6% for GPT-Image-2, without modifying the editor.
- SolveEdit has 2,728 goal-driven visual transformation cases.
- SolveScore measures completion and unintended changes without one reference image.
- The strongest evaluated model reaches only 57.0% SolveScore.
- SolveEdit-PLAN adds 9.1 points on average; GPT-Image-2 rises to 71.6%.
Full article185 words · extracted from huggingface.co · click to collapse
Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing that change without disturbing unrelated content. However, existing benchmarks mainly evaluate perception, generation, or explicitly specified transformations, leaving goal-driven visual problem solving underexplored. To bridge this gap, we introduce SolveEpIT, a benchmark for visual problem solving through scene transformation. Given an image and a goal, a model must infer a valid transformation from the request, the scene, or a visually expressed rule, then execute it while preserving unrelated content. SoLvEEDrr contains 2,728 cases. Atomic transition contracts specify required and protected conditions, enabling SoLvEScoRE to measure completion and unintended changes without a single reference output. The strongest evaluated model achieves only57.0% SolvEScore. We further introduce SolveEdiT-PLAN, a two-stage visual planner that instantiates the transition before generation. Under matched single-generation evaluation, it improves SoLvEScoRE by 9.1 points on average across three tested generators, including a gain from 57.0% to 71.6% for GPT-Image-2, without modifying the editor.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.35504