VibeEdit: Image Editing with Canvas Instructions
VibeEdit edits images from on-canvas marks, beating FireRed 79.9 to 67.4 on a 419-case benchmark.
VibeEdit lets users mark regions and optional notes directly on an image to add, remove, replace, modify, or move objects without a separate text prompt. The model adapts Qwen-Image-Edit with layer-decoupled conditioning and is trained on 1.55 million edit pairs using region-weighted fine-tuning and rubric-guided reinforcement learning. On a 419-case benchmark of similar objects, it scores 79.9 on a VLM rubric and 32.8 dB outside-region PSNR, versus 67.4 and 24.0 dB for FireRed.
- Users place spatial marks and short notes instead of writing a full prompt.
- Training uses 1.55 million masked source-target pairs with structured edit descriptions.
- VibeEdit adapts Qwen-Image-Edit with separate image and canvas conditioning.
- On 419 cases it scores 79.9 versus 67.4 for text-instructed FireRed.
Full article197 words · extracted from huggingface.co · click to collapse
In text-guided image editing, describing the desired change is often straightforward, but identifying the intended object or region can be cumbersome, especially when several objects look alike. We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image. Together, these annotations form a canvas instruction that specifies where to edit and what to change. Our editor, VibeEdit, follows these instructions to perform object addition, removal, replacement, attribute modification, and movement without a separate text prompt. We construct 1.55 million source-target edit pairs with object masks and structured edit descriptions, from which we render canvas instructions during training. We adapt Qwen-Image-Edit with layer-decoupled conditioning that separately encodes source images and canvas instructions for image editing. We train the model with region-weighted supervised fine-tuning, followed by rubric-guided reinforcement learning to improve edit completion, local edit quality, and preservation of unedited regions. We evaluate VibeEdit on an independently constructed, human-curated benchmark of 419 cases emphasizing target selection among similar objects. VibeEdit achieves a VLM rubric score of 79.9 and an outside-region PSNR of 32.8 dB, compared with 67.4 and 24.0 dB for FireRed, the highest-scoring text-instructed baseline in our evaluation.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.12229