Ambient @ EgoProactive 2026:基于视觉接地监督的主动式自我中心辅助
Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision
浏览论文内容
中文总结 AI 辅助
本研究提出单令牌分类与视觉接地监督方法,在ECCV 2026可穿戴AI挑战赛中获大模型组第一,显著提升干预决策性能。
中文摘要 AI 辅助
我们提交了参加ECCV 2026可穿戴AI挑战赛EgoProactive赛道的结果,在大模型组中排名第一,在<=2B参数组中排名第二。该任务要求可穿戴助手在每段八秒的自我中心视频后决定是否干预或保持沉默。我们的方法包含两个主要组成部分。首先,我们将干预时机重新表述为单令牌分类。模型不是生成$interrupt$<话语>或$silent$,而是预测是或否,我们从这两个令牌的重新归一化概率中得出决策。这种表述相比自由形式生成,宏F1提升了0.249,G均值提升了0.30。其次,由于标注数据仅限于发布的验证集,我们使用一个工具调用视频代理生成额外的监督,该代理检查每个片段并分配干预时间戳。仅叙述的替代方案规模大四倍且成本低十倍,但迁移效果不如来自无关真实语料库的监督,这表明对于该任务,视觉接地比标注数量更重要。
英文摘要
We present our submission to the EgoProactive track of the ECCV 2026 Wearable AI Challenge, which ranked first in the large-model division and second in the <=2B division. The task requires a wearable assistant to decide after each eight-second segment of egocentric video whether to intervene or remain silent. Our approach has two main components. First, we reformulate intervention timing as single-token classification. Rather than generating either $interrupt$<utterance> or $silent$, the model predicts yes or no, and we derive the decision from the renormalised probabilities of these two tokens. This formulation improved macro-F1 by 0.249 and G-mean by 0.30 over free-form generation. Second, because labelled data were limited to the released validation set, we generated additional supervision using a tool-calling video agent that inspects each clip and assigns intervention timestamps. A narration-only alternative was four times larger and ten times cheaper, but transferred worse than supervision from an unrelated real corpus, suggesting that visual grounding is more important than annotation volume for this task.
发表机构
- Team Ambient(Ambient团队)
机构由 AI 辅助整理,请以论文原文为准。