arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Ambient @ EgoLongQA 2026:将长视频感知蒸馏到亚20亿参数模型中

Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model

Logesh Kumar Umapathi

arXiv 2609.07154首次发表:更新:

发表机构

Team Ambient(环境团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

通过将智能体流水线的感知模块蒸馏到修剪后的2B模型中,以1.1%参数达到大型流水线89%准确率,在EgoLongQA挑战赛≤2B组别夺冠。

AI 中文摘要

我们描述了在ECCV 2026可穿戴AI挑战赛EgoLongQA赛道中的参赛方案,该方案在不超过20亿参数组别中排名第一,在留出测试集上得分为0.8279。我们的系统是一个单一的20亿参数视觉语言模型,通过一次贪婪前向传播回答关于十分钟第一人称视频的多项选择题;该模型通过将工具使用智能体流水线的初级感知模块(而非智能体本身)蒸馏到一个小型学生模型中,并使用过滤为回答正确的教师轨迹进行训练。它仅使用大型智能体流水线1.1%的参数,就达到了其89%的准确率。这使基础模型在我们的留出问题上从27.1%提升至81.4%。该20亿参数骨干网络有2.2132B参数,因此超过组别限制,为使参赛作品合规,我们将多语言嵌入表从248,320行修剪至143,469行,达到1.9985B参数,且在保留行上证明logits完全相同。

英文摘要

We describe our entry to the EgoLongQA track of the Wearable-AI Challenge in ECCV 2026, which placed first in the <=2B parameter division with 0.8279 on the held-out test set. Our system is a single 2B vision-language model that answers multiple-choice questions about ten-minute egocentric videos in one greedy forward pass; It is obtained by distilling the junior perception module of a tool-using agentic pipeline, not the agent itself into a small student, using teacher traces filtered to those that answered correctly. it reaches 89% of the accuracy of the large agentic pipeline using 1.1% of its parameters. This raises a 27.1% base model to 81.4% on our held-out questions. The 2B backbone has 2.2132B parameters and therefore over the divisional limit, to make the entry admissable we prune the multilingual embedding table from 248,320 to 143,469 rows, reaching 1.9985B with provably identical logits on retained rows.

CommentsWinning solution technical report for the EgoLongQA track of the ECCV 2026 Wearable AI Grand Challenge

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑