arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mudragen:几何监督生成用于保护印度古典舞蹈遗产的交互双手手印(Mudras)

Mudragen: Geometrically Supervised Generation of Interacting Two-Hand Mudras for Preserving Indian Classical Dance Heritage

Jagadish Kashinath Kamble, Jayanta Mukhopadhyay, Debaditya Roy, Partha Pratim Das

arXiv 2609.03415首次发表:更新:

发表机构

Indian Institute of Technology Kharagpur; Ashoka University(印度理工学院克勒格布尔分校; 阿育王大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对印度古典舞蹈低资源手势数据集及文本描述不足的问题,提出几何感知监督的MudraGen条件扩散框架,可生成真实协调的交互双手手印,性能优于现有模型,适用于文化保护与舞蹈教育。

AI 中文摘要

自动生成手势对于印度古典舞蹈的传承至关重要,也是其保护的关键。印度古典舞蹈手势数据集本质上是低资源的,且许多手印(mudras)的梵文规范定义缺乏精确的文本描述,限制了传统文本条件图像生成模型的有效性。我们提出了MudraGen,这是一个条件扩散框架,用于合成婆罗多舞(Bharatanatyam,一种印度古典舞蹈形式)中交互双手手势——合掌手印(Samyukta Hasta Mudras)的真实RGB图像。与之前针对简单手势或单手手势的研究不同,MudraGen引入了几何感知监督,以捕捉交互双手的精确协调性、解剖学有效性和文化细微差别。我们制定了三个几何感知目标:用于3D关节对齐的关键点损失、用于手间空间一致性的关节偏移损失,以及作为解剖学正则化器的形状一致性,该正则化器在允许双手姿态独立的同时鼓励手部形态一致。这些目标共同引导扩散模型生成解剖学上合理且协调良好的手部配置,从而能够合成逼真且姿态准确的手势图像。实验结果表明,MudraGen在视觉真实性、解剖学正确性和精细手部姿态结构的保留方面优于现有最先进的生成方法,能够忠实地再现复杂的合掌手印(Samyukta Hasta mudras)。除了定量增益外,其生成基于文化且结构一致的手势的能力,凸显了其在文化保护和舞蹈教育中的实际应用。

英文摘要

Automatic generation of hand gestures is essential for the transmission of Indian classical dance and critical for its preservation. Indian classical dance gesture datasets are inherently low-resource, and the canonical Sanskrit definitions of many mudras lack precise textual descriptions, limiting the effectiveness of conventional text-conditioned image generation models. We present \textbf{MudraGen}, a conditional diffusion framework that synthesizes realistic RGB images of \textit{Samyukta Hasta Mudras} -- interactive two-hand gestures from Bharatanatyam (an Indian classical dance form). Unlike prior work on simple hand signs or single-hand gestures, MudraGen introduces geometry-aware supervision to capture the precise coordination, anatomical validity, and cultural nuance of interacting hands. We formulate three geometry-aware objectives: Keypoint Loss for 3D joint alignment, Joint Offset Loss for inter-hand spatial coherence, and Shape Consistency, which serves as an anatomical regularizer by encouraging consistent hand morphology while allowing independent hand poses. Together, these objectives guide the diffusion model toward anatomically plausible and well-coordinated hand configurations, enabling the synthesis of photorealistic and pose-accurate gesture images. Experimental results show that MudraGen surpasses existing state-of-the-art generative approaches in visual realism, anatomical correctness, and preservation of fine hand-pose structure, enabling faithful reproduction of complex Samyukta Hasta mudras. Beyond quantitative gains, its ability to generate culturally grounded and structurally consistent gestures highlights practical applications in cultural preservation and dance education.

CommentsAccepted for publication in ACM Journal on Computing and Cultural Heritage (JOCCH) Special Issue on Visual Heritage

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑