arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越应答:建模人机声乐合奏中的互惠协调

Beyond Call and Response: Modelling Reciprocal Coordination in Human-AI Vocal Ensembles

Polina Proutskova

arXiv 2608.07376首次发表:更新:

发表机构

Industry Commons Foundation(产业公共基础基金会)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对无指挥声乐合奏的互惠协调问题,提出一种声乐智能体研究架构,探索人工歌手对人类音乐协作的影响。

AI 中文摘要

与AI的音乐互动常被组织为应答循环:人类表演,系统解读该动作,随后回应、伴奏或安排音乐事件。无指挥的声乐合奏则提出不同问题:歌手同时行动并持续相互影响,既无指挥、节拍器、伴奏、乐谱或调音源固定节拍与音高,集体组织源于多对多的互惠调整。本文将此类合奏建模为耦合动态系统,提出一种声乐智能体的研究架构,该架构会进入而非仅追踪集体状态。部分目标曲目有节拍,另一些则呈现无法简化为节拍网格的非等时时间轮廓,我们将后者视为通用框架的难题。该架构连接现场多通道采集、方言与歌唱感知表示、集体状态推理、声乐生成及原位评估。所提出的研究议程不仅关注人工歌手能否同步,更关注其存在如何重塑人类的协调、领导力、风格及音乐传播。

英文摘要

Musical interaction with AI is often organised as a response loop: a human performs, the system interprets that action, and the system answers, accompanies, or schedules a musical event. Unconducted vocal ensembles pose a different problem. Singers act simultaneously and continuously affect one another; neither timing nor pitch is fixed by a conductor, metronome, accompaniment, score, or tuning source. Collective organisation emerges from many-to-many reciprocal adjustment. This paper frames such ensembles as coupled dynamic systems and proposes a research architecture for vocal agents that enter, rather than merely track, their collective states. Some target repertoires are metrical, while others exhibit non-isochronous temporal contours that cannot be reduced to a beat grid; we treat the latter as a hard case for a general framework. The architecture connects multichannel capture in the field to dialect- and singing-aware representation, collective-state inference, vocal generation, and in-situ evaluation. The resulting agenda asks not only whether an artificial singer can synchronise, but how its presence reorganises human coordination, leadership, style, and musical transmission.

Comments5 pages, 2 figures

DOI:10.1145/3776591.3837051

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑