arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过鲁棒3D化身放置学习理解肢体语言

Learning to Understand Body Language from Flight through Robust 3D Avatar Placing

Dragos Costea, Alina Marcu, Cristina Lazar, Marius Leordeanu

arXiv 2607.27865首次发表:更新:

发表机构

National University of Science and Technology “Politehnica” Bucharest; “Simion Stoilow” Institute of Mathematics of the Romanian Academy; NORCE Norwegian Research Centre AS(布加勒斯特理工国立大学; 罗马尼亚科学院“西米翁·斯托伊洛夫”数学研究所; 挪威NORCE研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出Drones2BodyLanguage数据集,结合轻量几何世界模型训练12种架构,可提升空中机器人对人类交际意图的感知准确率,为社交智能无人机提供数据支撑与方法。

AI 中文摘要

远程感知人类动作与意图是具备社交智能的空中机器人的前提,但用于学习该能力的数据几乎不存在。我们推出Drones2BodyLanguage数据集,将人类动作与真实无人机镜头关联:把体现10种交际意图的化身置于未修改的4K无人机场景中,其位置、尺度和方向均符合度量标准,且在数百帧相机运动中保持稳定。实现这一目标的是本地场景的轻量几何世界模型——通过流单目深度将语义选定的锚点提升至3D空间,在该模型中,放置点被预测为具有可证刚性不变权重的仿射锚点组合,并在SVD拟合的地面旋转下重新渲染。在场景和动作不相交的划分上,针对12种架构,使用放置数据训练后,对于真实、重定向和生成的动作,平均意图准确率均大幅提升,且在两个野外场景中也验证了该增益。

英文摘要

Perceiving human motion and intent at long range is a prerequisite for socially intelligent aerial robots, yet the data to learn it barely exists. We introduce Drones2BodyLanguage, a dataset grounding human motion in real UAV footage: avatars manifesting ten communicative intents are placed into unmodified 4K drone scenes with metrically correct position, scale and orientation, maintained over hundreds of frames of camera motion. Enabling it is a lightweight geometric world model of the local scene - semantically selected anchors lifted to 3D through streaming monocular depth - in which a placement point is predicted as an affine anchor combination with provably rigid-invariant weights, and re-rendered under an SVD-fitted ground rotation. Across twelve architectures on scene- and motion-disjoint splits, training on placed data lifts mean intent accuracy by a wide margin for real, retargeted and generated motion alike, with gains confirmed on two in-the-wild scenes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑