arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40230cs.CVcs.AI

EviRover:强化智能体感知,超越一瞥之限

EviRover: Reinforcing Agentic Perception Beyond a Glance

Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang, Manyuan Zhang, Junke Wang, Xiangyu Yue

首次发表
浏览论文内容

中文总结 AI 辅助

EviRover通过智能体交互式感知解决证据不足问题,利用专门数据训练,在EviLens基准上超越骨干模型30分,性能媲美专有模型。

中文摘要 AI 辅助

视觉感知传统上被表述为对图像单次一瞥的一次性预测,其假设是图像内容和模型的参数化知识足以解决查询。这一假设在依赖细粒度视觉细节或需要知识密集型和最新信息的现实场景中往往失效。我们将此类情况称为“证据不足下的感知”,并将感知表述为一个能够获取超越单次一瞥信息的智能体过程。为解决该场景下数据缺失的问题,我们设计了两条专门的数据生成流水线,生成了用于训练的EviRover-SFT-5K和EviRover-RL-12K。我们进一步构建了EviLens,一个包含五个感知类别共688个实例的人工验证基准。基于这些数据,我们提出了EviRover,据我们所知,这是第一个通过交互明确训练以解决感知查询的感知智能体,采用监督微调后接智能体强化学习的方法。实验表明,4B规模的EviRover在EviLens上平均比其骨干模型高出30分,达到了与先进专有模型相当的性能。这些提升超越了EviLens,迁移至WebEyes、传统感知基准和通用多模态基准,包括在BrowseComp-VL上15分的提升。所有代码、模型和数据均已发布。

英文摘要

Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or require knowledge-intensive and up-to-date information. We term such cases \textit{perception under insufficient evidence} and formulate perception as an agentic process that can obtain information beyond a single glance. To address the absence of data for this setting, we design two dedicated data generation pipelines, yielding EviRover-SFT-5K and EviRover-RL-12K for training. We further construct EviLens, a human-verified benchmark comprising 688 instances across five perception categories. Building on these data, we present EviRover, to our knowledge the first perception agent explicitly trained to resolve perceptual queries through interaction, using supervised fine-tuning followed by agentic reinforcement learning. Experiments show that the 4B EviRover outperforms its backbone by 30 points on average on EviLens, reaching performance comparable to advanced proprietary models. The gains transfer beyond EviLens to WebEyes, conventional perception benchmarks, and general multimodal benchmarks, including a 15-point improvement on BrowseComp-VL. All code, models, and data are released.

发表机构

  • The Chinese University of Hong Kong(香港中文大学)
  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

↑