arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04589cs.CVcs.AI

2026年EgoVis首届EgoCross挑战赛:跨域自我中心视频问答

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering

Yuqian Fu, Tianwen Qian, Yanjun Li, Yu Li, Kunyu Peng, Xu Zheng, Yongqin Xian, Alessio Tonioni, Yanwei Fu, Xiaoling Wang, Danda Paudel, Federico Tombari, Luc Va… 展开作者

Yuqian Fu, Tianwen Qian, Yanjun Li, Yu Li, Kunyu Peng, Xu Zheng, Yongqin Xian, Alessio Tonioni, Yanwei Fu, Xiaoling Wang, Danda Paudel, Federico Tombari, Luc Van Gool, Leyi Wu, Yifan Zhao, Jinjie Zhang, Yinchuan Li, Yingcong Chen, Zixu Li, Zhiwei Chen, Zhiheng Fu, Wenbo Wang, Yupeng Hu, Weili Guan, Liqiang Nie, Takuya Murakawa, Toru Tamaki, Yi Wen, Zhenglin Du, Zhengyang Li, Lingling Li, Licheng Jiao, Wenping Ma

首次发表
浏览论文内容

中文总结 AI 辅助

本文介绍2026年EgoVis研讨会举办的首届EgoCross跨域自我中心视频问答挑战赛,含两个赛道,共获1500余份提交,公布结果与获胜方案,资源公开以推进相关研究。

中文摘要 AI 辅助

EgoCross是一项跨域自我中心视频问答基准,旨在评估多模态大语言模型是否能在普通日常生活场景之外实现泛化。首届EgoCross挑战赛在2026年CVPR的第三届EgoVis研讨会上举办,针对四个目标领域的第一人称视频对模型进行评估:外科手术、工业装配、极限运动及动物视角。每个测试样例包含一段自我中心视频片段、一个问题和四个候选答案,模型需从中选出正确选项。本技术报告介绍了挑战赛任务、基准资源及两个官方Codabench赛道。源数据受限赛道要求参与者仅使用官方基线模型和小型支持集,而开源赛道则允许在禁止手动构建目标域训练数据的规则下选择更广泛的模型和训练数据。本次挑战赛共收到来自130余名参与者的1500余份提交,其中开源赛道有19支队伍参与,源数据受限赛道有38支队伍参与。我们还公布了官方排行榜结果并总结了两个赛道的获胜方案。希望本报告能成为推进跨域自我中心视频理解的实用技术参考,所有资源包括挑战赛数据、基线实现及获胜队伍发布的代码均已公开。

英文摘要

EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scenarios. The first EgoCross Challenge was hosted at the Third EgoVis Workshop at CVPR 2026 and evaluated models on first-person videos from four target domains: surgery, industrial assembly, extreme sports, and animal perspectives. Each test example consists of an egocentric video clip, a question, and four candidate answers, from which the model must select the correct option. This technical report introduces the challenge task, benchmark resources, and two official Codabench tracks. The Source-Limited Track restricts participants to the official baseline model and a small support set, whereas the Open-Source Track permits broader choices of models and training data under rules that prohibit the manual construction of target-domain training data. In total, the challenge received more than 1,500 submissions from over 130 participants, with 19 teams participating in the Open-Source Track and 38 teams in the Source-Limited Track. We further present the official leaderboard results and summarize the winning solutions from both tracks. We hope that this report will serve as a useful technical reference for advancing cross-domain egocentric video understanding. All resources, including the challenge data, baseline implementation, and code released by the winning teams, are made publicly available.

补充信息

↑