arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MusiChat:用于音乐创作的氛围作曲

MusiChat: Vibe Composing for Music Creation

Callie C. Liao, Duoduo Liao, Ellie L. Zhang

arXiv 2607.24873首次发表:更新:

发表机构

Stanford University; George Mason University; IntelliSky(斯坦福大学; 乔治梅森大学; 智能天空)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对AI音乐创作中迭代难问题,提出MusiChat系统,通过分层可控框架、记忆增强架构及混合意图路由机制,实现人机协作音乐创作,经评估在交互准确率和用户喜好比例上表现良好,支持多轮音乐创作和人机共创。

AI 中文摘要

近期人工智能音乐生成的进展使用户能根据自然语言提示创作完整乐曲。但多数现有系统遵循提示-重新生成范式,迭代细化困难。我们提出MusiChat,一个对话式氛围作曲系统,通过自然语言交互和迭代细化实现人机协作音乐创作。其核心是分层可控音乐生成框架,分离歌词对齐的音乐结构生成与表达性表层实现。系统通过记忆增强架构集成大语言模型与混合符号音乐引擎,混合意图路由机制实现对精确音乐编辑和开放式创作请求的高效解释。通过客观分析和人类研究评估,单轮和多轮交互准确率分别达95.31%和100%,旋律自然度和音乐质量的喜好与不喜欢比例分别为2:1和3:1。结果表明MusiChat支持连贯的多轮音乐创作和人机交互共同创作。

英文摘要

Recent advances in AI music generation have enabled users to create complete musical pieces from natural-language prompts. However, most existing systems follow a prompt-and-regenerate paradigm, making iterative refinement difficult because users must repeatedly recreate compositions instead of directly evolving existing musical ideas. We present MusiChat, a conversational vibe composing system that enables collaborative human-AI music creation through natural-language interaction and iterative refinement. At the core of MusiChat is a hierarchical controllable music generation framework that separates lyric-aligned musical structure generation from expressive surface realization, allowing flexible stylistic transformations and structure-preserving edits. The system integrates a large language model with a hybrid symbolic music engine through a memory-augmented architecture that maintains the active composition state and user history across interactions. A hybrid intent-routing mechanism further enables efficient interpretation of both precise musical edits and open-ended creative requests. Rather than regenerating compositions from scratch, MusiChat incrementally transforms an evolving musical artifact while preserving relevant musical structure and user intent. We evaluate MusiChat through objective analysis and human studies, achieving 95.31% and 100% accuracy for single- and multi-turn interactions, respectively, and obtaining like-to-dislike ratios of 2:1 for melody naturalness and 3:1 for musical quality. Our results demonstrate that MusiChat supports coherent multi-turn music authoring and interactive human-AI co-creation through a conversational interface.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑