arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20245cs.CV

SAGE-Yoga:用于瑜伽姿势分类和关节级纠正的多线索学习

SAGE-Yoga: Multi-Cue Learning for Yoga Pose Classification and Joint-Level Correction

Hung Le Chi, Khanh Minh Huynh, Long Nghia Tran Pham, Tan Phuc Huynh, Trong-Thuan Nguyen, Minh-Triet Tran

首次发表
浏览论文内容

中文总结 AI 辅助

SAGE-Yoga提出统一粗到细框架,结合互补视觉集成与选择性几何验证,实现瑜伽姿势分类和关节级纠正,在Yoga-82上达到90.7%Top-1准确率。

中文摘要 AI 辅助

自动瑜伽分析需要准确的姿势分类和对姿势执行的可解释反馈。然而,现有方法通常依赖于单一的视觉预测,难以区分视觉上相似的姿势,并将姿势分类和纠正视为独立任务。为解决这些局限性,我们提出了SAGE-Yoga,一个从单张RGB图像进行瑜伽姿势分类和关节级纠正的统一粗到细框架。受瑜伽教练使用多种互补线索评估姿势的启发,SAGE-Yoga首先采用基于装袋的互补视觉骨干集成来生成候选姿势类别的排序集合。此外,基于边距的门控机制保留高置信度的视觉预测,仅对模糊案例进行几何验证。再者,一旦确定最终姿势类别,SAGE-Yoga检索一个中值参考姿势,并将观察到的关节角度与类别特定分布进行比较,以识别未对齐的关节。最后,这些偏差被转化为可操作的纠正反馈。实验上,在Yoga-82数据集上的实验表明,视觉集成达到89.0%的Top-1准确率,而完整框架将性能提升至90.7%的Top-1准确率和90.1%的Macro-F1。这些结果表明,将互补的视觉证据与选择性几何验证相结合,可改进细粒度姿势分类,同时实现可解释的关节级纠正。

英文摘要

Automated yoga analysis requires both accurate pose classification and interpretable feedback on pose execution. However, existing methods often rely on a single visual prediction, struggle to distinguish visually similar poses, and treat pose classification and correction as separate tasks. To address these limitations, we propose SAGE-Yoga, a unified coarse-to-fine framework for yoga pose classification and joint-level correction from a single RGB image. Inspired by how yoga instructors assess posture using multiple complementary cues, SAGE-Yoga first employs a bagging-based ensemble of complementary visual backbones to generate a ranked set of candidate pose classes. Additionally, a margin-based gating mechanism preserves confident visual predictions while invoking geometric verification only for ambiguous cases. Moreover, once the final pose class is determined, SAGE-Yoga retrieves a medoid reference pose and compares the observed joint angles with class-specific distributions to identify misaligned joints. Finally, these deviations are translated into actionable corrective feedback. Empirically, experiments on the Yoga-82 dataset show that the visual ensemble achieves 89.0% Top-1 accuracy, while the complete framework improves performance to 90.7% Top-1 accuracy and 90.1% Macro-F1. These results demonstrate that combining complementary visual evidence with selective geometric verification improves fine-grained pose classification while enabling interpretable, joint-level correction.

补充信息

↑