AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation
AutoRef automatically rewrites image-generation harnesses, lifting FLUX.2 klein 4B from 5.72 to 7.37 on MultiBanana.
AutoRef automatically optimizes the executable harness for agentic multi-reference image generation while keeping the generator and reasoning models frozen. A coding agent iteratively rewrites harness code, using separate tasks for proposing changes and selecting candidates and continuing from a beam of top-ranked harnesses. The resulting AutoRef-Harness raised open-weight FLUX.2 klein 4B from 5.72 to 7.37 on held-out four-reference MultiBanana tasks, matching or exceeding Nano Banana Pro and GPT-Image-1.5. Without re-optimization, the same harness also improved results when the generator, reference count, benchmark, evaluator, or reasoning model changed.
- A coding agent rewrites the harness while both models stay frozen.
- Proposal feedback is separated from tasks used to select candidates.
- FLUX.2 klein 4B improved from 5.72 to 7.37 on held-out MultiBanana tasks.
- The harness transferred across generators, benchmarks, evaluators, and reasoning models.
Full article279 words · extracted from huggingface.co · click to collapse
Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally pasted. Recent work has proposed image generation agents that combine image generation models, reasoning models, and a harness, which is an executable program that specifies how reference images are interpreted, how generation is performed, how outputs are diagnosed, and how the final image is selected. In multi-reference generation, however, references play different roles and outputs must satisfy many criteria at once, such as fidelity to each reference and the naturalness of the whole image, so many parts of the harness could be improved, from how references are processed to how outputs are diagnosed. This makes it hard to predict which changes will improve performance and by how much, and good harnesses difficult to design by hand; indeed, human-written harnesses vary widely in performance. We therefore propose AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code. AutoRef separates the tasks whose feedback informs proposals from the tasks used to select candidates, and continues the search from a beam of the top-ranked harnesses on the selection tasks. Using this procedure, we discover AutoRef-Harness, which improves the open-weight FLUX.2 [klein] 4B from 5.72 to 7.37 on held-out four-reference tasks of the MultiBanana benchmark, matching or exceeding proprietary models including Nano Banana Pro and GPT-Image-1.5. Without re-optimization, the same harness also improves results when the generator, number of references, benchmark, evaluator, or reasoning model differs from those used in the search.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.35530