arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38086cs.CV

VISTA:通过同策略蒸馏内化集体视觉经验用于主动多模态智能体

VISTA: Internalizing Collective Visual Experience via On-Policy Distillation for Active Multimodal Agents

Zheng Jiang, Houde Qian, Yiming Chen, Ling Li, Chaoyang Li, Yueqi Li, Yuxuan Liu, Lifeng Sun

首次发表
浏览论文内容

中文总结 AI 辅助

VISTA通过同策略蒸馏将同一输入的多轨迹视觉观察转化为共享监督,内化集体经验,提升主动多模态智能体的感知与推理性能。

中文摘要 AI 辅助

主动多模态智能体在推理过程中使用视觉工具获取与任务相关的证据。尽管强化学习对每个输入采样多条交互轨迹,基于结果的优化目标主要利用群体来估计标量优势,导致互补的视觉发现未被充分利用。我们提出VISTA,通过同策略蒸馏将来自同一输入轨迹的观察转化为共享监督,从而内化集体视觉经验。集体视觉经验蒸馏(CVED)将这些观察与其交互上下文组织起来,并将其与个体决策对齐,而异质性感知策略改进(HAPI)强化成功轨迹,并为失败尝试提供经验引导的蒸馏。经验条件教师评估学生采样的响应前缀,允许一条轨迹中的发现指导另一条轨迹的学习,而无需替换学生原有的历史或生成新的目标轨迹。训练后的智能体保留其视觉工具,并使用自身的交互历史进行行动。VISTA在评估的同等规模主动多模态智能体中取得了最强的平均性能,并在细粒度感知和一般推理任务上持续优于同骨干训练基线,证明了集体经验对主动多模态学习的价值。

英文摘要

Active multimodal agents use visual tools to acquire task-relevant evidence while reasoning. Although reinforcement learning samples multiple interaction trajectories per input, outcome-based objectives primarily use the group to estimate scalar advantages, leaving complementary visual discoveries underused. We introduce VISTA, which internalizes collective visual experience through on-policy distillation by turning observations from same-input rollouts into shared supervision. Collective visual experience distillation (CVED) organizes these observations with their interaction context and aligns them with individual decisions, while heterogeneity-aware policy improvement (HAPI) reinforces successful trajectories and provides experience-guided distillation for unsuccessful attempts. An experience-conditioned teacher evaluates the student's sampled response prefixes, allowing discoveries from one trajectory to guide learning in another without replacing the student's original history or generating new target trajectories. The trained agent retains its visual tools and acts using its own interaction history. VISTA achieves the strongest average performance among the evaluated active multimodal agents of comparable size and consistently outperforms same-backbone training baselines across fine-grained perception and general reasoning tasks, demonstrating the value of collective experience for active multimodal learning.

发表机构

  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑