DeepVoyager-VL
Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
1Peking University
2HKUST (GZ)
3NUDT
4OUC
5HITSZ
6Huawei Cloud BUVision in the loop, not only at the input.
The supervision gap. As Figure 1(a)–(c) shows, existing methods build the reasoning chain in text first and place vision at the input or near the answer, so acquired images rarely drive later retrieval.
EventVoyage-VL is our structure-before-language synthesis pipeline and resulting vision-in-the-loop data source. As Figure 1(d)–(e) shows, it creates long-horizon questions with intermediate visual dependencies; the curated trajectories are then distilled into DeepVoyager-VL.
Comparison of multimodal search data synthesis paradigms

Figure 1. Comparison of multimodal search data synthesis paradigms. Prior methods place vision at the input through entity substitution (a), concentrate visual reasoning before text-based search (b), or graft visual evidence near the answer (c). We instead synthesize vision-in-the-loop questions from a visually enriched multimodal event graph (d), as illustrated by a representative long-horizon example (e).
Overview of DeepVoyager-VL

Figure 2. Overview of DeepVoyager-VL, encompassing vision-in-the-loop data synthesis, difficulty-aware trajectory curation, and supervised agent training.
Leading multimodal search performance at both scales.
Across ten multimodal search benchmarks, DeepVoyager-VL-8B and DeepVoyager-VL-30B-A3B outperform scale-matched multimodal search agents on eight and nine benchmarks, respectively, while the same framework also transfers consistently across three different base models.
Performance across ten multimodal search benchmarks
Accuracy or task-native score; higher is better.
| Model | MMSearch | SimpleVQA | LiveVQA | FVQA | BC-VL | MM-BC | MMSearch+ | VDR | BC-V³ | VisBrowse | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Direct Answer | |||||||||||
| GPT-5.5 | 68.7 | 67.0 | 73.0 | 66.7 | 47.9 | 17.5 | 20.3 | 18.6 | 23.0 | 36.1 | 43.9 |
| Gemini-3.1-Pro | 64.2 | 64.1 | 65.0 | 58.9 | 41.4 | 11.5 | 26.4 | 15.6 | 19.3 | 23.7 | 39.0 |
| Claude-Opus-4.6 | 59.8 | 71.7 | 53.1 | 60.1 | 43.5 | 13.2 | 13.2 | 15.4 | 15.0 | 27.2 | 37.2 |
| Qwen3-VL-30B-A3B-Instruct | 18.7 | 53.2 | 42.7 | 34.7 | 29.6 | 4.0 | 3.2 | 3.8 | 6.0 | 11.2 | 20.7 |
| Qwen3-VL-8B-Instruct | 15.2 | 44.7 | 41.0 | 28.0 | 25.1 | 4.9 | 3.2 | 2.8 | 1.0 | 8.9 | 17.5 |
| Agentic Workflow | |||||||||||
| GPT-5.5 | 82.7 | 82.3 | 90.3 | 84.3 | 68.2 | 51.9 | 48.6 | 42.0 | 55.0 | 66.3 | 67.2 |
| Gemini-3.1-Pro | 82.7 | 81.0 | 87.3 | 81.3 | 65.9 | 44.4 | 51.5 | 40.8 | 50.0 | 62.1 | 64.7 |
| Claude-Opus-4.6 | 81.7 | 81.7 | 88.0 | 79.3 | 63.2 | 49.4 | 52.1 | 36.6 | 55.0 | 62.7 | 65.0 |
| Qwen3-VL-30B-A3B-Instruct | 64.7 | 71.0 | 73.3 | 72.3 | 41.6 | 9.9 | 17.7 | 21.6 | 11.3 | 23.1 | 40.7 |
| Qwen3-VL-8B-Instruct | 61.7 | 60.0 | 67.7 | 69.7 | 34.6 | 5.8 | 13.5 | 16.8 | 7.7 | 15.4 | 35.3 |
| Multimodal Deep Search Agents · 8B | |||||||||||
| MMSearch-R1-7B | 53.8 | 57.4 | 48.4 | 58.4 | — | — | — | — | 4.0 | — | — |
| WebWatcher-7B | 49.1 | 54.3 | 51.2 | — | 21.2 | — | — | — | 4.7 | — | — |
| DeepEyesV2-7B | 63.7 | 59.4 | — | 60.6 | — | — | — | — | — | — | — |
| SenseNova-MARS-8B | 67.8 | 70.2 | 56.2 | 67.1 | — | — | — | — | — | — | — |
| Vision-DeepResearch-8B | 69.6 | — | 76.7 | 64.7 | 42.6 | — | 20.4 | 29.2 | 11.7 | — | — |
| MM-DeepResearch-8B | 67.8 | 65.9 | 65.0 | 69.2 | 37.9 | — | — | — | — | — | — |
| POINTS-Seeker-8B | 70.8 | 68.8 | 77.7 | 71.2 | 44.4 | — | 25.2 | — | — | — | — |
| OpenSearch-VL-8B | 64.5 | 71.6 | 59.6 | 71.5 | 37.6 | — | — | 20.8 | — | — | — |
| SimpleSearch-VL-8B | 77.1 | 76.6 | 75.2 | 76.8 | 52.1 | — | 32.5 | — | — | — | — |
| Visual-Seeker-8B | 72.2 | — | — | — | 47.6 | 16.1 | 27.3 | — | — | 34.7 | — |
| DeepVoyager-VL-8B | 72.7 | 76.3 | 82.7 | 82.7 | 58.4 | 24.0 | 37.1 | 35.0 | 32.3 | 47.3 | 54.8 |
| Multimodal Deep Search Agents · 30B | |||||||||||
| WebWatcher-32B | 55.3 | 59.0 | 58.7 | — | 27.0 | — | — | — | 8.7 | — | — |
| SenseNova-MARS-32B | 74.3 | 74.1 | 60.8 | 72.6 | — | — | — | — | — | — | — |
| Skywork-R1V4-30B-A3B | 66.1 | — | — | 67.2 | 38.4 | — | — | — | — | — | — |
| Vision-DeepResearch-30B-A3B | 69.6 | — | 77.6 | 74.2 | 53.7 | — | 28.5 | 37.8 | — | — | — |
| REDSearcher-MM-30B-A3B † | 72.9 | — | 79.3 | — | 57.2 | 23.5 | 26.6 | — | — | — | — |
| MM-DeepResearch-32B | 69.0 | 67.6 | 68.0 | 70.1 | 43.0 | — | — | — | — | — | — |
| LMM-Searcher-30B-A3B †‡ | 71.0 / 72.3 | — | — | — | — | 22.3 / 30.1 | 32.9 / 34.8 | — | — | 42.0 / 48.3 | — |
| OpenSearch-VL-30B-A3B | 68.7 | 74.9 | 67.4 | 73.2 | 41.1 | — | — | 33.5 | — | — | — |
| SimpleSearch-VL-30B-A3B | 83.6 | 79.6 | 81.1 | 79.0 | 55.9 | — | 34.4 | — | — | — | — |
| DeepVoyager-VL-30B-A3B | 74.0 | 81.0 | 82.7 | 84.7 | 64.2 | 30.5 | 40.6 | 39.4 | 35.0 | 53.8 | 58.6 |
Framework comparison across base models
All framework evaluations use up to 50 interaction turns.
| Base model | Evaluation method | MMSearch | MM-BC | MMSearch+ | VisBrowse | Avg. |
|---|---|---|---|---|---|---|
| GPT-5 | Direct Answer | 33.3 | 10.3 | 19.1 | 26.0 | 22.2 |
| Vision-DeepResearch workflow | 63.7 | — | 17.2 | — | — | |
| LMM-Searcher agentic search | 72.2 | 23.7 | 34.8 | 35.5 | 41.6 | |
| DeepVoyager-VL (Ours) | 79.7 | 49.8 | 50.8 | 59.2 | 59.8 | |
| Gemini-2.5-Pro | Direct Answer | 39.8 | 10.3 | 14.5 | 27.2 | 23.0 |
| Vision-DeepResearch workflow | 69.0 | — | 22.2 | — | — | |
| LMM-Searcher agentic search | 66.3 | 12.1 | 28.1 | 16.0 | 30.6 | |
| DeepVoyager-VL (Ours) | 72.0 | 19.0 | 33.4 | 43.8 | 42.1 | |
| Qwen3-VL-30B-A3B-Thinking | Direct Answer | 17.7 | 7.1 | 2.7 | 13.0 | 10.1 |
| Vision-DeepResearch workflow | 53.2 | — | 13.6 | — | — | |
| LMM-Searcher agentic search | 62.0 | 9.8 | 14.4 | 16.0 | 25.6 | |
| DeepVoyager-VL (Ours) | 70.0 | 10.0 | 20.9 | 20.7 | 30.4 |
Both the data and the agent design matter.
Adding 7K vision-in-the-loop trajectories yields further average gains of +5.4 for 8B and +6.5 for 30B, while removing any core agent component lowers average performance by 2.5–4.0 points.
Training-data ablation
| Setting | BC-VL | BC-V³ | VisBrowse | Avg. |
|---|---|---|---|---|
| Qwen3-VL-8B (Agentic) | 34.6 | 7.7 | 15.4 | 19.2 |
| + 20K multi-source trajectories | 54.4 | 26.7 | 40.8 | 40.6 |
| + 7K VIL trajectories | 58.4 | 32.3 | 47.3 | 46.0 |
| Qwen3-VL-30B-A3B (Agentic) | 41.6 | 11.3 | 23.1 | 25.3 |
| + 20K multi-source trajectories | 62.2 | 28.7 | 42.6 | 44.5 |
| + 7K VIL trajectories | 64.2 | 35.0 | 53.8 | 51.0 |
Framework component ablation
| Setting | FVQA | VDR | BC-V³ | Avg. |
|---|---|---|---|---|
| Full DeepVoyager-VL | 78.0 | 36.0 | 39.3 | 51.1 |
| w/o Summary | 74.0 | 35.0 | 32.3 | 47.1 |
| w/o Image Search | 77.7 | 35.4 | 32.7 | 48.6 |
| w/o Fetch Image | 76.3 | 35.6 | 33.3 | 48.4 |
| w/o Crop Image | 75.7 | 33.0 | 36.3 | 48.3 |
The learned behavior is longer and more visual.
Under unified rollouts, EventVoyage-VL yields 64.3% visual-tool calls and peaks at 16–20 turns, compared with at most 40.6% visual-tool use and 1–10-turn peaks in prior datasets.
of tool calls are visual, versus 40.6%, 30.3%, and 10.1% in three public datasets.
average points from 7K VIL trajectories on the 30B model, beyond a strong 20K mixture.
achieves 54.8 and 58.6 average scores without an additional reinforcement-learning stage.
Trajectory analysis
(a) Visual tool engagement
(b) Interaction horizonFigure 3. Trajectory statistics from unified Doubao-2.0-Pro rollouts on 1,000 examples per dataset. (a) Tool-call proportions by functional category. (b) Trajectory counts in five-turn intervals. OSVL, VDR, and RED-MM denote OpenSearch-VL, Vision-DeepResearch, and REDSearcher-MM, respectively.
Cite DeepVoyager-VL.
@misc{zhang2026deepvoyagervlincentivizingvisionintheloopsearch,
title={DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents},
author={Huanyao Zhang and Jiepeng Zhou and Runhao Zhao and Yanzhe Shan and Jiaoyang Chen and Bowen Zhou and Bo Li and Fang Wang and Jialong Wu and Zhengwei Tao and Lang Mei and Xiaohan Yu and Liyan Liu and Chong Chen and Wentao Zhang},
year={2026},
eprint={2608.01827},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.01827},
}