发表机构
University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PAMoR将情感转化为可量化的V-A控制参数,结合动作与情感先验,在Unitree G1机器人上实时生成可编辑的情感动作,其情感识别率接近人类表演水平且优于基线。
AI 中文摘要
在社交场景中,人们读取人形机器人的动作时,不仅关注其执行的动作,还关注其传递的情感。目前,携带情感的动作生成仅针对人类化身,其风格来自参考片段或情感词汇,而这两者都无法进行量化参数化。我们提出了PAMoR,它将情感转化为可测量的控制参数:通过机器人运动学原生计算的效价-唤醒度(V-A)坐标。该坐标可通过姿态扩展和运动能量以闭式形式获得,这些测量值直接作为生成条件,无需人工标注。在共享潜在空间中训练的动作先验和两个情感先验,在每个去噪步骤中进行组合:动作先验确定执行的内容,情感先验调节执行的方式。全身动作在具有29个自由度的Unitree G1机器人上实时自回归生成,且动作和情感均可编辑。生成的动作在其全范围内跟踪命令的V-A值,同时文本到动作的保真度仍与仅文本基线相匹配。在一项感知研究中,评估者在38%的试验中识别出命令的情感,高于两个基线,接近人类表演身体的44%识别率。
英文摘要
People read a humanoid robot's motion in social settings not only for the action performed but for the affect conveyed. Motion carrying that affect has so far been generated for human avatars, where style is taken from a reference clip or an emotion word, neither of which can be quantitatively parameterized. We present PAMoR, which turns affect into a measured control parameter: a valence-arousal (V-A) coordinate computed natively on robot kinematics. It is obtained in closed form from postural expansion and movement energy, and these measurements serve directly as generation conditions, with no human annotation. An action prior and two affect priors, trained in a shared latent space, are composed at each denoising step: the action prior fixes what is performed, the affect priors modulate how. Whole-body motion rolls out autoregressively on a 29-DoF Unitree G1 in real time, with action and affect both editable. Generated motion tracks the commanded V-A over its full range while text-to-motion fidelity still matches text-only baselines. In a perceptual study, raters identify the commanded emotion on 0.38 of trials, above both baselines and approaching the 0.44 reported for acted human bodies.
CommentsUnder Review of IEEE Robotics and Automation Letters, 8 pages