发表机构
Saarland University; Max Planck Institute for Informatics; IIT Delhi(萨尔兰大学; 马克斯·普朗克信息学研究所; 印度理工学院德里分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过探测运动分词器码本,发现重建训练虽能解码几何和利手性,但难以捕获手势语义,提示需基于语义属性评估以改进分词器设计。
AI 中文摘要
离散运动分词器将运动编码为原子单元,并广泛用于共语手势生成。目前尚不清楚从这些码本中可恢复哪些运动属性,尤其是与手势语义相关的属性。我们使用19个共语手势描述符(涵盖从原始运动到抽象交际功能)对一个基于重建训练的分词器码本进行了探测。结果表明,几何特征和利手性可以从词元嵌入中轻松解码,而运动类别虽在离散码使用上表现出系统性差异,却仅能被弱解码。重建质量与描述符可解码性之间的这种差距表明,仅靠重建目标并不能保证手势语义被捕获,而基于此类属性评估码本可以指导更具语义性的运动分词器的设计。
英文摘要
Discrete motion tokenizers encode motion as atomic units and are widely used for co-speech gesture generation. It remains unclear which motion properties, especially those relevant to gesture semantics, are recoverable from these codebooks. We probe a reconstruction-trained codebook using 19 co-speech gesture descriptors spanning from raw motion to abstract communicative function. Results show that geometry and handedness are readily decodable from token embeddings, while motion category is only weakly decoded despite showing systematic differences in discrete code usage. This gap between reconstruction quality and descriptor decodability suggests that reconstruction objectives alone do not guarantee that gesture semantics are captured, and that evaluating codebooks on such properties can guide the design of more semantic motion tokenizers.
CommentsAccepted at MINT EMNLP 2026