arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.04903cs.CV

InterSing:用于3D二重唱动画及更广泛应用的显式交互动力学框架

InterSing: Explicit Interaction Dynamics for 3D Duet Singing Animation and Beyond

  • School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院)
  • Guangdong Engineering Center for Large Model and GenAI Technology(广东省大模型与生成式人工智能技术工程中心)
  • State Key Laboratory of Subtropical Building and Urban Science(亚热带建筑与城市科学国家重点实验室)
  • Ministry of Education Key Laboratory of Big Data and Intelligent Robot(大数据与智能机器人教育部重点实验室)
  • School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算与信息系统学院)

机构由 AI 辅助整理,请以论文原文为准。

Yihan Zhou, Zikai Huang, Yuyang Yu, Xuemiao Xu, Cheng Xu, Shengfeng He

AI总结:

该研究提出InterSing框架,通过引入交互logits建模二重唱的显式交互动力学,生成协调且逼真的3D演唱头部动画,可推广至多表演者表演并实现直观控制。

AI中文摘要:

我们提出了InterSing,一个用于生成逼真3D头部动画的框架,适用于二重唱表演。与独唱不同,二重唱表演要求每位歌手在展现个人表现力的同时,在音乐上的关键节点(如乐句边界、同步节奏和问答式段落)进行间歇性互动。由于这些互动是稀疏且依赖节奏的,现有的音频驱动动画方法和对话交互模型无法充分捕捉其结构。我们的核心见解是,二重唱的协调可表示为随时间变化的信号,反映表演者在整首歌曲中相互参与的程度。基于这一观察,我们引入了交互logits,这是一种可解释的潜在表示,用于建模每个时间步中跨表演者参与的程度。我们使用弱监督学习这些logits,并将其用于条件交互感知扩散模型,该模型由音频特征和交互动力学共同驱动。这种表述支持统一的多模态生成,涵盖独立动作、协调行为以及它们之间的平滑过渡。实验表明,InterSing生成的逼真且富有表现力的演唱头部动画,比现有方法具有更强的协调性和音乐对齐性,同时保留每位表演者的特征动作风格。我们进一步证明,相同的表述可推广至多表演者表演,并能直观控制表演者何时以及如何参与。

英文摘要:

We present InterSing, a framework for generating realistic 3D head animations for duet singing performances. Unlike solo singing, duet performance requires each singer to balance individual expressiveness with intermittent interaction at musically salient moments, such as phrase boundaries, synchronized rhythms, and call-and-response passages. Because these interactions are sparse and rhythm-dependent, existing audio-driven animation methods and conversational interaction models do not adequately capture their structure. Our key insight is that duet coordination can be represented as a time-varying signal that reflects how strongly performers engage with one another throughout a song. Based on this observation, we introduce interaction logits, an interpretable latent representation that models the degree of cross-performer engagement at each time step. We learn these logits using weak supervision and use them to condition an interaction-aware diffusion model jointly driven by audio features and interaction dynamics. This formulation enables unified multi-mode generation, spanning independent motion, coordinated behavior, and smooth transitions between them. Experiments show that InterSing generates realistic and expressive singing head animations with stronger coordination and musical alignment than existing methods, while preserving each performer's characteristic motion style. We further demonstrate that the same formulation generalizes to multi-singer performances and provides intuitive control over when and how performers engage.

↑