arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06972cs.CV

BoT-Feedback:基于生物力学证据的多模态推理用于可解释人体动作反馈

BoT-Feedback: Grounding Multimodal Reasoning in Biomechanical Evidence for Explainable Human Action Feedback

Xu Dong, Wanqing Li, Anthony Adeyemi-Ejeye, Andrew Gilbert

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模态大语言模型在人体动作反馈中产生泛泛建议和幻觉的问题,BoT-Feedback通过四阶段生物力学推理框架,将推理锚定在结构化生物力学证据上,显著提升反馈质量、可解释性和鲁棒性,平均专家评分提升40%。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)在视觉理解和多模态推理方面展示了令人印象深刻的能力,但在人体动作反馈生成方面仍然存在根本性的局限。现有方法直接从视觉观察中推断指导反馈,产生泛泛的建议、有限的可解释性以及物理上不合理的幻觉。相比之下,专家级人类教练通过关节运动学、姿态和身体动力学的显式生物力学推理来诊断表现。我们引入了BoT-Feedback,一个将MLLM推理锚定在结构化生物力学证据上的框架。我们的关键贡献是“生物力学思维”(BoT),一个四阶段推理框架,逐步识别动作、定位关键身体区域、分析专家和学生表现之间的定量生物力学差异,并综合生成可解释的指导反馈。为支持这一推理过程,我们开发了一个即插即用的生物力学数据解析器(BDP),将视频转换为结构化的生物力学描述符,以及一种对齐策略,在时间上匹配专家和学生的动作。我们进一步引入了BiomAF,一个包含成对师生视频、3D骨架、生物力学属性和专家指导标注的基准。在十二个开源和闭源MLLM上的实验表明,将推理锚定在生物力学证据上持续提高了反馈质量、可解释性和鲁棒性,同时大幅减少了生物力学幻觉。BoT-Feedback将平均专家评估分数从2.07提高到2.95(+40%),使紧凑的开源MLLM能够接近规模大得多的专有系统在可解释动作反馈生成方面的性能。

英文摘要

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in visual understanding and multimodal reasoning, yet they remain fundamentally limited in Human Action Feedback Generation. Existing methods infer coaching feedback directly from visual observations, producing generic advice, limited interpretability, and physically implausible hallucinations. In contrast, expert human coaches diagnose performance through explicit biomechanical reasoning over joint kinematics, posture, and body dynamics. We introduce BoT-Feedback, a framework that grounds MLLM reasoning in structured biomechanical evidence. Our key contribution is Biomechanics of Thought (BoT), a four-stage reasoning framework that progressively identifies the action, localises the critical body regions, analyses quantitative biomechanical differences between expert and student performances, and synthesises interpretable coaching feedback. To support this reasoning process, we develop a plug-and-play Biomechanical Data Parser (BDP) that converts videos into structured biomechanical descriptors and an alignment strategy that temporally matches expert and student motions. We further introduce BiomAF, a benchmark containing paired teacher-student videos, 3D skeletons, biomechanical attributes, and expert-coaching annotations. Experiments across twelve open- and closed-source MLLMs demonstrate that grounding reasoning in biomechanical evidence consistently improves feedback quality, interpretability, and robustness while substantially reducing biomechanical hallucinations. BoT-Feedback improves the average expert evaluation score from 2.07 to 2.95 (+40%), enabling compact open-source MLLMs to approach the performance of substantially larger proprietary systems for explainable action feedback generation.

发表机构

  • University of Surrey(萨里大学)
  • University of Wollongong(伍伦贡大学)

机构由 AI 辅助整理,请以论文原文为准。

↑