OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
OneSearch-VL-8B beats tool-using Qwen3-VL-8B by up to 20.2 points on new visual research benchmarks.
OneSearch-VL is a unified multimodal deep-research agent built around a Visually Grounded Evidence Graph that tracks links among localized visual anchors, entity relations, source-backed facts, and answer operations. A VGEG data engine produces OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K, and an Evidence-aware Visual-Grounded Rubric reward supervises RL. OneSearch-VL-8B improves over tool-using Qwen3-VL-8B by 20.2 and 17.6 percentage points on new multi-image and video benchmarks, with further gains on seven image benchmarks and VideoDR.
- VGEG links visual anchors, relations, sources, and answer operations.
- Training sets are OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K.
- EVGR reward supervises evidence traceability and visual grounding.
- OneSearch-VL-8B gains 20.2 and 17.6 points over tool-using Qwen3-VL-8B.
Full article175 words · extracted from huggingface.co · click to collapse
Single-image, multi-image, and video deep research require different visual operations but share a workflow of visual grounding, external retrieval, and fact composition. A key challenge is to preserve the dependencies linking localized visual anchors, entity relations, source-supported facts, and answer-producing operations. We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation. Our VGEG-based data engine constructs and verifies multi-image and video questions and filters expert trajectories. Using these data, we assemble OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K for SFT and RL, respectively. We further derive the Evidence-aware Visual-Grounded Rubric reward (EVGR) from VGEG annotations to supervise evidence traceability and visual grounding during RL. For fine-grained evaluation, we construct OneSearch-MI-Bench and OneSearch-Video-Bench, organizing questions by the research operations encoded in their VGEGs. Experiments show that OneSearch-VL-8B improves over Qwen3-VL-8B with tool access by 20.2 and 17.6 percentage points on the two new benchmarks, respectively, while also achieving substantial gains across 7 image benchmarks and VideoDR. Project repository: https://github.com/appletea233/OneSearch-VL
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.12419