语言增强的视频动作预测:设计基础、基准与开放挑战
Language-Augmented Video Action Anticipation: Design Fundamentals, Benchmarks, and Open Challenges
浏览论文内容
中文总结 AI 辅助
本综述提出证据感知设计图谱,按任务类型与语言信息介入点组织视频动作预测文献,并审计Ego4D-LTA与EPIC-KITCHENS-100基准,将LLM收益等结论表述为需匹配验证的假设。
中文摘要 AI 辅助
动作预测旨在从不完整的视频和不确定的时间上下文中预测未来的动作。近期系统在不同阶段引入大型语言模型(LLM)、视觉-语言模型(VLM)或语言派生语义,但当任务形式、视觉预训练、监督方式、解码器设计和评估代码同时变化时,所报告的性能提升难以解释。本综述的核心贡献是一个基于证据的设计图谱,该图谱将任务类型与语言派生信息介入的点进行交叉。我们从六个轴描述任务类型,这些轴将文献组织为五大类任务族:单动作、序列、物体交互、跨视角和规划导向设置。C1-C3将干预定位在上下文构建、目标/意图建模和未来解码中,而C4被视为相邻的、新兴的接地/可执行性扩展。与通用处理流水线不同,该图谱将每次干预与适当的反事实、失败诊断和可允许的证据声明联系起来。辅助贡献包括对Ego4D-LTA和EPIC-KITCHENS-100的协议级审计、多维证据概况以及骨干感知比较与消融协议(BCAP)。未解决的EK-100记录被视为报告可比性案例研究,而非排行榜。因此,关于LLM收益、目标模糊性和视野效应的证据被表述为需要匹配验证的可测试假设,而非因果结论。随附包包含本综述中使用的编码证据、源定位器、协议元数据和版本化目录。
英文摘要
Action anticipation predicts future human actions from partial video under incomplete context and temporal uncertainty. Recent systems introduce large language models (LLMs), vision-language models (VLMs), or language-derived semantics at different stages, but reported gains are difficult to interpret when task formulation, visual pretraining, supervision, decoder design, and evaluation code change simultaneously. The central contribution of this review is an evidence-aware design map that crosses task regime with the point at which language-derived information intervenes. We characterise task regimes along six axes. These axes organise the literature into five broad task families: single-action, sequence, object-interaction, cross-view, and planning-oriented settings. C1-C3 locate interventions in context construction, goal/intention modelling, and future decoding, while C4 is treated as an adjacent, emerging grounding/executability extension. Unlike a generic processing pipeline, the map links each intervention to an appropriate counterfactual, failure diagnosis, and permissible evidence claim. Supporting contributions include a protocol-level audit of Ego4D-LTA and EPIC-KITCHENS-100, a multidimensional evidence profile, and the Backbone-Aware Comparison and Ablation Protocol (BCAP). The unresolved EK-100 record is treated as a reporting-comparability case study and is not used as a leaderboard. Evidence for LLM benefits, goal ambiguity, and horizon effects is therefore formulated as testable hypotheses requiring matched validation, not as causal conclusions. The accompanying package contains the coded evidence, source locators, protocol metadata, and versioned catalogue used in the review.