局部视频理解能否跨遭遇迁移?EgoGears基准
Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark
浏览论文内容
中文总结 AI 辅助
针对具身系统跨遭遇的局部视频理解迁移问题,提出EgoGears基准,含单视频和多视频问题,发现所有模型在多视频任务上平均下降22.5个百分点,瓶颈在于观察-证据绑定与路由状态跟踪。
中文摘要 AI 辅助
具身系统必须使在一次遭遇中获得的知识在另一次遭遇中可用,尽管视点、运动和光照会发生变化。然而,跨视频的总体准确率混淆了局部感知的失败与保持观察身份、建立对应关系以及组合证据的失败,从而掩盖了局部视频理解是否真正迁移。我们引入EgoGears,一个互补的单视频和多视频基准,旨在诊断这一转变。它包含567个单视频和1487个多视频问题,这些问题源自126段人工收集的自我中心记录,覆盖39条户外路线。在不同运动速度和光照条件下的重复穿越使比较基于共享的物理环境;531个问题需要跨独立记录进行对齐。单视频问题衡量模型可用的局部视觉、空间和运动证据,而多视频问题则测试证据是否仍绑定到正确的观察,并能组合成一致的路由关系。我们在主排行榜中报告了六个模型家族的29个单视频和31个多视频MLLM配置。在两种分割上可比评估的20个配置中,每个模型在多视频问题上的表现都更差,平均下降22.5个百分点,并且在固定答案格式和评分时,这一差距仍然存在。这一差距不能简单地由额外视频或记录边界来解释。核心瓶颈是观察-证据绑定和有序路由状态跟踪。代码和基准在此https URL公开可用。
英文摘要
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video benchmark designed to diagnose this transition. It contains 567 single-video and 1,487 multi-video questions derived from 126 human-collected egocentric recordings covering 39 outdoor routes. Repeated traversals across movement speeds and lighting conditions ground comparisons in shared physical environments; 531 questions require alignment across independent recordings. Single-video questions measure the local visual, spatial, and motion evidence available to a model, while multi-video questions test whether evidence remains bound to the correct observation and can be composed into consistent route relationships. We report 29 single-video and 31 multi-video MLLM configurations across six model families in the main leaderboard. Among the 20 configurations evaluated comparably on both splits, every model performs worse on multi-video questions, with a mean decrease of 22.5 percentage points, and the gap persists when answer format and scoring are held fixed. The gap is not explained simply by additional videos or recording boundaries. The central bottlenecks are observation--evidence binding and ordered route-state tracking. The code and benchmark are publicly available at https://github.com/lei-qi-233/EgoGears.
发表机构
- INSAIT, Sofia University “St. Kliment Ohridski”(索非亚大学“圣克利门特·奥赫里德斯基”INSAIT)
- University of Würzburg(维尔茨堡大学)
- Technical University of Munich(慕尼黑工业大学)
- Institute of Information Engineering(信息工程研究所)
- Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
机构由 AI 辅助整理,请以论文原文为准。