面向跨数据集自我中心凝视建模的基于运动的分词
Motion-Based Tokenization for Cross-Dataset Egocentric Gaze Modeling
浏览论文内容
中文总结 AI 辅助
该研究针对跨数据集自我中心凝视建模,提出基于运动的分词方法,通过多维度评估对比不同表示方式,发现其在部分迁移场景下的优势,明确事件构建对跨数据集迁移的影响。
中文摘要 AI 辅助
凝视正日益被用作视觉模型和多模态模型的输入信号,但目前尚未就如何在不同数据集间表示凝视达成共识。原始轨迹保留了细节,但存在噪声且依赖设备;而粗略的事件标签易于建模,却可能丢失局部运动结构。我们将事件对齐、固定时间范围的角位移定义为一种可解释的、以事件为条件的运动词汇,并将其与仅基于事件、基于空间、绝对角度、学习到的向量量化(VQ)以及连续表示进行比较。为了评估迁移能力、目标可预测性以及分词崩溃情况,我们的评估结合了下一个分词预测、目标域遗憾、低阶目标参考、配对自助法、顺序敏感性、基序重叠以及冻结结构探针。在一个事件对齐的头戴式设备基准测试中,角运动分词在一种迁移方向上的目标域遗憾低于冻结码本VQ分词,而在反向迁移方向上则无定论。探针揭示了互补的表示属性,仅基于事件的分词显示,低困惑度可能几乎无法保留运动信息。在第三个自我中心数据集上,对I-VT、原生事件和帧跨度接口的匹配比较表明,事件构建会显著改变迁移效果:原生事件迁移至EGTEA数据集时遗憾最低,而帧跨度事件的基序重叠为零,作为源数据时表现极差。因此,基于运动的分词为事件对齐的自我中心凝视流提供了一种紧凑的表示,同时该评估明确了目标可预测性和事件构建如何影响跨数据集结论。
英文摘要
Gaze is increasingly used as an input signal for vision and multimodal models, yet no consensus exists on how to represent it across datasets. Raw traces preserve detail but are noisy and device-dependent, while coarse event labels are easy to model but can discard local motion structure. We formulate event-aligned, fixed-horizon angular displacement as an interpretable, event-conditioned motion vocabulary and compare it with event-only, spatial, absolute-angle, learned vector-quantized, and continuous representations. To assess transfer alongside target predictability and token collapse, our evaluation combines next-token prediction with target-domain regret, low-order target references, paired bootstrap, order sensitivity, motif overlap, and frozen structural probes. In an event-aligned headset benchmark, angular-motion tokens have lower target-domain regret than frozen-codebook VQ tokens in one transfer direction, while the reverse direction is inconclusive. The probes reveal complementary representation properties, and event-only tokens show that low perplexity can retain little motion information. On a third egocentric dataset, a matched comparison of I-VT, native, and frame-span interfaces shows that event construction materially changes transfer: native events have the lowest regret into EGTEA, while frame-span events have zero motif overlap and fail severely as a source. Motion-based tokenization therefore provides a compact representation for event-aligned egocentric gaze streams, while the evaluation identifies how target predictability and event construction shape cross-dataset conclusions.
发表机构
- Technical University of Munich (TUM)(慕尼黑工业大学(TUM))
- Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
机构由 AI 辅助整理,请以论文原文为准。