发表机构
University of Exeter(埃克塞特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对可穿戴动作预测在时间损坏下的可靠性问题,提出结合时间可靠性抑制与鲁棒动词-名词图解码的轻量框架,显著提升损坏准确率与鲁棒性并减少稀有组合预测。
AI 中文摘要
可穿戴动作预测系统必须在缺失帧、掩蔽和传感器噪声的情况下保持可靠性,然而现有的自我中心动作预测方法大多假设观测数据是干净的。我们识别出时间损坏下的两种互补失效模式:编码期间不可靠的时间证据,以及解码期间不合理、低支持的动词-名词组合。我们通过一个轻量级框架来解决这些问题,该框架结合了时间可靠性抑制(TRS)和鲁棒动词-名词图(RVG)解码。TRS从投影的输入嵌入中预测每帧的抑制分数,并将其用作每个编码器块中学习到的键侧注意力惩罚,同时用于推导可靠性加权的时间池化。RVG利用从训练标签构建的基于PMI的兼容性图对动词-名词对进行重新排序。在损坏增强训练下,TRS+CA+RVG在六种损坏(包括训练期间未出现的三种机制)中达到了29.1%的平均损坏准确率和88.2%的相对鲁棒性,同时将稀有动词-名词预测从15.4%降低到1.1%。多种子和诊断实验表明,TRS对合成掩蔽有响应;打乱图和仅频率对照进一步表明,RVG的收益依赖于真正的成对兼容性,而非仅边际频率效应。
英文摘要
Wearable action anticipation systems must remain reliable despite missing frames, masking, and sensor noise, yet existing egocentric anticipation methods largely assume clean observations. We identify two complementary failure modes under temporal corruption: unreliable temporal evidence during encoding and implausible, low-support verb-noun compositions during decoding. We address them with a lightweight framework combining Temporal Reliability Suppression (TRS) and Robust Verb-Noun Graph (RVG) decoding. TRS predicts a per-frame suppression score from the projected input embedding and uses it as a learned key-side attention penalty at every encoder block and to derive reliability-weighted temporal pooling. RVG re-ranks verb-noun pairs using a PMI-based compatibility graph constructed from training labels. Under corruption-augmented training, TRS+CA+RVG reaches 29.1% average corrupted accuracy and 88.2% relative robustness across six corruptions, including three mechanisms absent during training, while reducing rare verb-noun predictions from 15.4% to 1.1%. Multi-seed and diagnostic experiments show that TRS responds to synthetic masking; shuffled-graph and frequency-only controls further indicate that RVG gains depend on genuine pairwise compatibility rather than marginal-frequency effects alone.
Comments16 pages, 4 figures. Accepted at the ECCV 2026 Workshop on Towards Real-time Multimodal Contextual Assistants (Wearable AI Workshop)