发表机构
University of Tehran(德黑兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对双手动作分割中跨手时间延迟问题,提出LACA轻量模块,结合空状态抑制无效信息传递,集成到Polyphony后在HA-ViD和ATTACH数据集上提升了分割性能,还推出无未来信息的变体LACA-C。
AI 中文摘要
双手动作分割通常在相同时间索引处融合左右手表示,不过协调的手部转换可能存在非零且随时间变化的延迟。我们提出感知延迟跨手对齐(Lag-Aware Cross-Hand Alignment, LACA),这是一个轻量型模块,可显式估计特定手部特征流之间的定向时间偏移分布。LACA从估计的偏移中提取跨手信息,并结合学习到的空状态,在不支持兼容跨手转换时抑制信息传递。对齐使用从帧级训练标注自动推导的兼容感知目标进行监督,无需额外标签。对HA-ViD和ATTACH训练标注的分析显示,分别有44.7%和48.9%的转换锚点存在稳健的非零跨手匹配,而时间偏移对照组的这一比例仅为18.6%和21.3%。将LACA集成到Polyphony中后,与我们复现的Polyphony基线相比,在HA-ViD上双手平均F1@50从40.4提升至42.5,边界F1从56.5提升至59.6;在ATTACH上分别从19.9提升至21.8,从44.7提升至47.9。这些提升仅需约0.0086百万个额外可训练参数。我们进一步提出LACA-C,这是一种无未来信息的变体,将对齐和完整推理流程限制在当前及过去观测中。在ATTACH上,LACA-C实现了83.6%的转换线索召回率、种子平均中位可用性延迟233毫秒、每分钟0.72个错误线索,以及分割阶段每秒224.9个当前位置预测的吞吐量。这些结果表明,显式跨手时间对齐可同时提升动作分割和边界定位性能,且支持及时的无未来感知。
英文摘要
Dual-hand action segmentation commonly fuses left- and right-hand representations at identical temporal indices, although coordinated hand transitions may occur with nonzero and time-varying delays. We introduce Lag-Aware Cross-Hand Alignment (LACA), a lightweight module that explicitly estimates directional temporal-offset distributions between hand-specific feature streams. LACA retrieves cross-hand information from the estimated offsets and incorporates a learned null state to suppress transfer when no compatible cross-hand transition is supported. Alignment is supervised using compatibility-aware targets derived automatically from frame-level training annotations, without requiring additional labels. Analysis of the HA-ViD and ATTACH training annotations reveals robust nonzero cross-hand matches for 44.7% and 48.9% of transition anchors, respectively, compared with 18.6% and 21.3% under temporally shifted controls. When integrated into Polyphony, LACA improves the two-hand mean F1@50 from 40.4 to 42.5 and boundary F1 from 56.5 to 59.6 on HA-ViD, and from 19.9 to 21.8 and 44.7 to 47.9, respectively, on ATTACH, relative to our reproduced Polyphony baseline. These gains require only approximately 0.0086 million additional trainable parameters. We further introduce LACA-C, a future-free variant that restricts alignment and the complete inference pipeline to current and past observations. On ATTACH, LACA-C achieves 83.6% transition-cue recall, a seed-averaged median availability delay of 233~ms, 0.72 false cues per minute, and segmentation-stage throughput of 224.9 current-position predictions per second. These results demonstrate that explicit cross-hand temporal alignment improves both action segmentation and boundary localization while supporting timely future-free perception.