返回全部动态

DeepVoyager-VL:激励长时程多模态智能体的视觉在环搜索

原标题:DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

Hugging Face Daily Papers一手来源研究质量 81

AI 摘要

DeepVoyager-VL 是一个用于长时程多模态智能体的视觉在环搜索框架,通过构建多模态事件图驱动数据合成,设计主动视觉获取和按需图像加载的智能体框架,并在合成数据上微调模型,无需强化学习。在十个多模态搜索基准上的实验证明了其有效性,推动了多模态深度研究智能体的发展。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents Abstract Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for open-world information access, evolving from single-turn factual retrieval toward l


发布时间:—
抓取时间:2026-08-04 15:17
来源机构:Hugging Face
阅读原文huggingface.co