arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15683cs.CV

V-ICAL Bench:评估交互环境中多模态智能体的视频上下文学习

V-ICAL Bench: Evaluating Video In-Context Learning for Multimodal Agents in Interactive Environments

Ziqian Fan, Shibo Xu, Junjie Li, Xiangyu Zhao, Shengyuan Ding, Yifan Yang, Zhenjie Yang, Haodong Duan, Yue Zhou, Zhihang Zhong, Xue Yang

首次发表
浏览论文内容

中文总结 AI 辅助

V-ICAL基准通过342个交互任务评估多模态智能体的视频上下文学习,发现最佳模型仅得54.4分,远低于人类83.6分,揭示了关键能力差距。

中文摘要 AI 辅助

虽然上下文学习(ICL)使模型能够在不更新参数的情况下从示例中适应,但多模态上下文学习在很大程度上仍未得到充分探索,特别是在交互环境中的视频演示方面。对于多模态智能体而言,从视频中学习提出了独特的挑战:它们必须将上下文演示转化为可执行的策略,将这些策略应用于新的视觉状态,并根据环境反馈迭代地改进动作。我们引入了V-ICAL,这是一个新颖的基准,旨在评估多模态智能体的基于视频的上下文学习。V-ICAL包含37个环境中的342个交互式任务,利用人工策划的演示视频作为任务特定的行为示例,通过从目标初始化的持续交互来评估智能体。该基准将上下文知识归纳与核心智能体能力无缝连接,包括状态基础、时间记忆、规划和动态环境中的适应。对19个最先进的多模态智能体的广泛评估揭示了显著的局限性:表现最佳的模型Seed-2.1-Pro仅获得54.4/100的分数,而其他领先模型(例如Gemini-3.1-Pro、GPT-5.6)未能超过50,远低于人类基线83.6。对照比较进一步表明,当前智能体难以可靠地将视频示例转化为有效策略,未能产生一致的性能提升。最终,V-ICAL暴露了多模态智能体ICL能力中的关键差距,强调了对未来研究的迫切需求。

英文摘要

While In-Context Learning (ICL) enables models to adapt from exemplars without parameter updates, multimodal ICL remains largely underexplored, particularly regarding video demonstrations in interactive environments. For multimodal agents, learning from videos presents unique challenges: they must translate in-context demonstrations into executable policies, ground these policies in novel visual states, and iteratively refine actions based on environmental feedback. We introduce V-ICAL, a novel benchmark designed to evaluate video-based ICL for multimodal agents. Comprising 342 interactive tasks across 37 environments, V-ICAL utilizes human-curated demonstration videos as task-specific behavioral exemplars, evaluating agents through sustained interaction from a target initialization. The benchmark seamlessly connects in-context knowledge induction with core agentic capabilities, including state grounding, temporal memory, planning, and adaptation in dynamic environments. Extensive evaluations across 19 state-of-the-art multimodal agents reveal significant limitations: the best-performing model, Seed-2.1-Pro, achieves a score of only 54.4/100, while other leading models (e.g., Gemini-3.1-Pro, GPT-5.6) fail to surpass 50, far below the human baseline of 83.6. Controlled comparisons further demonstrate that current agents struggle to reliably translate video exemplars into effective policies, failing to yield consistent performance gains. Ultimately, V-ICAL exposes a critical gap in the ICL capabilities of multimodal agents, underscoring an urgent need for future research.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • South China University of Technology(华南理工大学)
  • Fudan University(复旦大学)
  • Microsoft Research Asia(微软亚洲研究院)
  • The University of Hong Kong(香港大学)
  • The Chinese University of Hong Kong(香港中文大学)
  • East China Normal University(华东师范大学)

机构由 AI 辅助整理,请以论文原文为准。

↑