arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于测量人机交互预测中跨数据集线索迁移的固定参考姿态残差

Fixed-Reference Pose Residuals for Measuring Cross-Dataset Cue Transfer in Human-Robot Interaction Anticipation

Bowen Yang, Xinliang Xiao, Wenjing Zhang, Li Yang, Wei Zhou

arXiv 2610.12245首次发表:更新:

发表机构

School of Automation, Nanjing University of Science and Technology(南京理工大学自动化学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出固定参考姿态残差(FRPR)模型,发现人机交互预测中跨数据集线索迁移存在不对称性,该模型用于测量而非提升预测精度,头部朝向残差可小幅提升双向性能。

AI 中文摘要

公共空间中的社交与服务机器人需要预测附近哪个人即将接近并触碰自己,以便在接触前做好响应准备。目前尚不清楚,当用一个机器人训练的模型在另一地点的另一个机器人上使用时,哪些线索能支持这种预测。我们用固定参考姿态残差(FRPR)模型研究该问题:该模型是基于人体边界框和掩码构建的几何预测器,经训练后被冻结,随后时间网络从人体姿态中学习其对数几率的加性修正,使每个预测恰好分解为几何项和姿态项。在由不同机器人记录的两个公共 egocentric 数据集 HUI360 和 SSUP-A 之间,所有选择均基于源数据,姿态修正将平均精度(AP)从 0.277 提升至 0.321(从 SSUP-A 到 HUI360),而在相反方向未产生可测量增益;在更强的源选择几何参考上也存在相同的不对称性。冻结操作未比联合训练带来 AP 优势,且简单几何基线和树集成仍具竞争力或更优,因此该结构用于测量而非预测精度。头部朝向残差在两个方向均带来小幅增益。事后分析显示,人是否面向相机的判别方向在数据集间保持一致,而头部俯仰角则发生反转。在源数据选择的阈值下,使用几何的神经模型最多检测到目标交互的 17%。代码和处理后的数据可在该 https URL 获取。

英文摘要

Social and service robots in public spaces need to anticipate which nearby person is about to approach and touch them, so that a response can be prepared before contact. It is largely unknown which cues support this anticipation when a model trained with one robot is used on another robot at a different site. We study this question with a fixed-reference pose residual (FRPR) model: a geometry predictor built from the person's bounding box and mask is trained and frozen, and a temporal network then learns from body pose an additive correction to its logit, so that every prediction splits exactly into a geometry term and a pose term. Between two public egocentric datasets recorded by different robots, HUI360 and SSUP-A, with every choice made on source data, the pose correction raised average precision (AP) from 0.277 to 0.321 from SSUP-A to HUI360 and gave no measurable gain in the opposite direction; the same asymmetry held over a stronger, source-selected geometry reference. Freezing gave no AP advantage over joint training, and simple geometric baselines and tree ensembles remained competitive or better, so the construction serves measurement rather than prediction accuracy. A head-orientation residual added small gains in both directions. Post hoc, whether a person faces the camera kept its discriminative direction across datasets, whereas head pitch reversed. With thresholds chosen on source data, the neural models that use geometry detected at most 17% of target interactions. Code and processed data are available at https://github.com/WeiZhou96/FRPR-interaction-anticipation.

Comments20 pages, 5 figures, 1 table. Supplementary material available as an ancillary file

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑