AI 中文总结
研究针对运动技能指导中专业教练稀缺的问题,提出AIDE框架,仅训练阶段用专业参考、推理阶段仅用学习者姿态生成反馈,在ExpertAF数据集上性能优于无参考基线,与需双阶段专业演示的方法相当。
AI 中文摘要
生成运动技能的自然语言指导反馈可加速学习,但专业教练稀缺且成本高昂。现有基于参考的方法在训练和推理阶段均需专业演示,限制了实际部署。我们提出AIDE(Automated Instruction via Distilled Expertise,基于提炼专业知识的自动化指令),该框架仅在训练阶段利用专业参考,推理阶段仅从学习者的姿态序列生成反馈。教师模型首先通过冻结的语言模型从配对的学习者-专家姿态中学习生成反馈,生成分别编码学习者姿态和学习者-专家差异的学习者标记与差异标记。学生模型继承教师的编码器和权重初始化,用辅助模块替代显式专家比较,仅从学习者姿态生成互补标记。在ExpertAF数据集上,AIDE在多数指标上优于无参考基线,性能与训练和推理阶段均需专业演示的方法相当,基于大语言模型(LLM)的评估也支持这些发现。
英文摘要
Generating natural-language coaching feedback on motor skills can accelerate learning, yet expert coaches are scarce and expensive. Existing reference-based methods require expert demonstrations at both training and inference time, limiting practical deployment. We propose AIDE (Automated Instruction via Distilled Expertise), a framework that exploits expert references only during training and generates feedback from a learner's pose sequence alone at inference. A teacher model first learns to generate feedback from paired learner-expert poses via a frozen language model, producing separate learner tokens and difference tokens that encode the learner-expert difference. A student model then inherits the teacher's encoder and weight initialization, replacing the explicit expert comparison with an auxiliary module that produces complementary tokens from the learner's pose alone. On the ExpertAF dataset, AIDE outperforms reference-free baselines on most metrics and performs comparably to methods requiring expert demonstrations at both training and inference, with LLM-based evaluation supporting these findings.
CommentsAccepted to ACM Multimedia 2026. 12 pages (including supplementary material), 4 figures