发表机构
Peter L. Reichertz Institute for Medical Informatics, Hannover Medical School; Lower Saxony Center for Artificial Intelligence and Causal Methods in Medicine (CAIMed)(汉诺威医学院彼得·L·赖歇茨医学信息学研究所; 下萨克森州医学人工智能与因果方法中心(CAIMed))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MacJEPA通过掩码上下文查询JEPA和窗口局部模态丢弃,在未修剪自我中心视频中实现视听识别对时间局部传感器中断的鲁棒性,单阶段联合优化,无需测试时适应,在Epic-Kitchens-100和Epic-Sounds上超越缺失模态基线。
AI 中文摘要
视听模型通过利用互补线索改进了自我中心动作识别,但通常假设推理时两个流都可用。现有的缺失模态方法在修剪过的单事件片段上操作,其中流完全存在或缺失,而真实传感器在长时间未修剪的观察中会失效并恢复。我们将自我中心模态缺失重新定义为未修剪多事件观察中时间局部化的传感器中断,以整个片段缺失作为极限情况。我们引入了MacJEPA,一种缺失模态鲁棒的掩码上下文查询JEPA,它通过提供的间隔查询在视听上下文上识别视觉动作和声学事件。窗口局部模态丢弃在训练期间模拟这些传感器中断。MacJEPA进一步将JEPA中的掩码从自监督借口重新利用为监督鲁棒性目标,对齐多模态内容标记和任务条件查询的掩码和干净潜在表示。所有目标与识别在单阶段中联合优化,无需测试时适应。在Epic-Kitchens-100和Epic-Sounds上,单个检查点在完整输入下保持竞争力,并且在移除主导或辅助流时始终超过已发布的缺失模态基线。因此,MacJEPA将强完整输入识别与时间缺失模态鲁棒性统一在一个操作于未修剪多事件视频的单一模型中。
英文摘要
Audio-visual models improve egocentric action recognition by exploiting complementary cues, yet typically assume that both streams remain available at inference. Existing missing-modality methods operate on trimmed, single-event clips in which a stream is entirely present or absent, whereas real sensors fail and recover within long, untrimmed observations. We redefine egocentric modality missingness as temporally localized sensor outages within untrimmed, multi-event observations, with whole-clip absence as the limiting case. We introduce \textbf{MacJEPA}, a missing-modality-robust \textbf{Ma}sked-\textbf{c}ontext query \textbf{JEPA} that recognizes visual actions and acoustic events from supplied interval queries over audio-visual context. Window-local modality dropout simulates these sensor outages during training. MacJEPA further repurposes masking in JEPA from a self-supervised pretext into a supervised robustness objective, aligning masked and clean latent representations of both multimodal content tokens and the task-conditioned queries. All objectives are optimized jointly with recognition in a single stage, requiring no test-time adaptation. Across Epic-Kitchens-100 and Epic-Sounds, a single checkpoint remains competitive under complete input and consistently surpasses published missing-modality baselines when either the dominant or auxiliary stream is removed. MacJEPA thus unifies strong full-input recognition with temporal missing-modality robustness in a single model operating on untrimmed multi-event videos.