arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25561cs.CL

EgoArgus:将视觉语言模型(VLM)作为模态接地用户支持情境助手的基准测试

EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports

Yu-Chien Tang, Yu-Hsiang Liu, An-Zi Yen

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对当前VLMs作为日常助手的模态仲裁难题,构建人工标注数据集EgoArgus,发现现有VLMs仍难胜任该任务,且现有偏差缓解方法效果有限,为部署提供了参考。

中文摘要 AI 辅助

视觉语言模型(VLMs)日益被定位为能感知第一人称环境、遵循用户对话并决定如何提供帮助的日常助手。现有以自我为中心的基准主要孤立评估视觉理解,留下了一个待解决的问题:当视觉证据与用户提供的语言处于有用、无关或冲突的情况时,模型能否在二者之间进行仲裁。我们推出EgoArgus,这是一个人工标注的数据集,用于在五个对话视频日常场景中评估以自我为中心的助手的理解与决策任务。我们的结果表明,将当前VLMs作为可靠的以自我为中心的助手仍具挑战性,这需要识别哪种模态更值得信任,并确定何时需要干预。更深入的分析还显示,现有的模态偏差缓解方法在提升性能方面相当受限,这为从业者将当前VLMs部署为日常助手提供了见解。

英文摘要

VLMs are increasingly positioned as daily assistants that perceive first-person environments, follow user dialogue, and decide how to help. Existing egocentric benchmarks mainly evaluate visual understanding in isolation, leaving open whether models can arbitrate between visual evidence and user-provided language when the two are helpful, irrelevant, or conflicting. We introduce EgoArgus, a human-annotated dataset for evaluating egocentric assistants on understanding and decision tasks in five dialogue-video daily scenarios. Our results demonstrate that it is still challenging for current VLMs as reliable egocentric assistants, which requires identifying which modality is trustworthy and deciding when intervention is warranted. Deeper analysis also shows that existing modality bias mitigation methods are quite restricted to enhance performance, providing insights to aid practioners into the deployment of current VLMs as daily assistants.

发表机构

  • National Yang Ming Chiao Tung University(国立阳明交通大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑