arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22519cs.HC

编舞基因组:将文本的静默结构放大为舞蹈

The Choreographic Genome: Amplifying the Silent Structure of Text into Dance

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Michael Li, Alison Ding

AI总结:

提出一种具身可视化工具,将文本的字节级结构而非语义映射为舞蹈动作,通过动作码本和物理平滑生成全身运动,展示被忽略的结构并追问信号与噪声的边界。

AI中文摘要:

生成式人工智能的最新进展使得以空前保真度合成复杂人体运动成为可能。然而,当前的文本到动作系统严格依赖语言语义:如果输入文本为“我举起双手”,模型会搜索一个双手举起的姿势,而文本的所有非语义属性都被当作噪声丢弃。在本工作中,我们将这种被丢弃的结构视为信号。我们提出了一种具身可视化工具,它放大的不是文本的含义,而是文本的构建方式。我们的方法首先利用主成分分析和K-Means聚类,将舞蹈运动学量化到一个包含256种风格“区域”的动作码本中,并沿运动的主轴对这些区域进行排序。然后,我们将任意输入文本的原始字节表示直接映射到该码本上,生成一个确定性的区域序列,我们称之为文本的“编舞基因组”。一个预计算的合理性图以及一组物理平滑例程将该基因组转化为流畅的全身运动,从而使舞蹈身体成为语义系统所忽略的字节级结构的展示表面。通过一系列艺术案例研究,包括一首莎士比亚十四行诗、一份机器错误日志、源代码、一位废奴主义者的提问,以及原住民和天城文文本,我们展示了每种文本都会产生视觉上截然不同的舞蹈,并且被以ASCII为中心的计算所边缘化的文本,每个字符所产生的动作量几乎是其近三倍。我们不将此视为一个动作合成基准,而是一种批判性和诗意的可视化,它追问我们选择将什么视为信号,以及我们允许什么被忽略。

英文摘要:

Recent advances in generative artificial intelligence have enabled the synthesis of complex human motion with unprecedented fidelity. However, current text-to-motion systems rely strictly on linguistic semantics: if an input reads "I put my hands up", the model searches for a pose with raised hands, and every non-semantic property of the text is discarded as noise. In this work, we treat that discarded structure as the signal. We present an embodied visualization instrument that amplifies not what a text means, but how it is built. Our method first quantizes dance kinematics into a motion codebook of 256 stylistic "regions" using Principal Component Analysis and K-Means clustering, and orders those regions along the dominant axis of movement. We then map the raw byte representation of any input text directly onto this codebook, producing a deterministic sequence of regions that we call the text's "choreographic genome". A precomputed plausibility graph and a set of physics smoothing routines turn this genome into fluid, full-body movement, so that the dancing body becomes a display surface for the byte-level structure that semantic systems ignore. Through a series of artistic case studies, including a Shakespeare sonnet, a machine error log, source code, an abolitionist's question, and Indigenous and Devanagari scripts, we show that each text produces a visibly distinct dance, and that scripts marginalized by ASCII-centric computing are amplified into close to three times as much movement per character. We frame this not as a motion-synthesis benchmark, but as a critical and poetic visualization that asks what we choose to count as signal, and what we allow to go unheard.

补充信息

↑