arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10360cs.HCcs.AIeess.AS

MazzikaAI:一种基于知识的性能到提示编译器,用于结合流式文本到音乐模型的实时阿拉伯maqam伴奏

MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model

  • Grand Valley State University(大峡谷州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Jiaxin Du, Boulbaba Abdeljaouad, Yong Zhuang, Haoyu Li

AI总结:

本文提出基于知识的MazzikaAI系统,通过编译实时输入生成文本提示,引导未微调的Google Lyria RealTime实现实时阿拉伯maqam伴奏,可精准控制微分音生成,为文化包容性生成音频提供了可扩展范式。

AI中文摘要:

阿拉伯maqam音乐具有微分音、调式特征,且建立在带装饰的呼应结构之上,是生成式音乐模型服务不足的传统音乐类型之一,这类模型的训练框架仍以西方音乐和平均律为主。实时伴奏进一步拉大了这一差距:AI伴奏伙伴必须具备聆听、动态适应的能力,且要遵循地道的微分音结构。流式文本到音乐模型具备强大的生成能力,但缺乏精确的控制接口。本文提出MazzikaAI,这是一种基于知识的系统,利用自然语言作为实时控制环路的执行器,通过将实时MIDI、手势和推断的和声编译为持续更新的文本提示,MazzikaAI在不微调模型的情况下引导未修改的流式生成器Google Lyria RealTime。该系统嵌入了6种核心maqam、典型装饰音和合奏动态的专业知识,保持实时响应性,按键到可听更新延迟小于1秒。实证评估表明,动态提示编译能可靠地将生成内容锚定在微分音音阶中,与基线生成相比,显著增加了非网格四分音内容。除核心实现外,MazzikaAI还展示了基于确定性知识的规则如何有效连接专业知识、非西方音乐传统和未微调的基础模型,该架构为实时人机协同创作建立了可扩展范式,为交互式伴奏、自适应音乐教育以及全球各类文化语境下的文化包容性生成音频提供了可推广蓝图。

英文摘要:

Arabic maqam music microtonal, modal, and built on ornamented call and response is among the traditions most underserved by generative music models, whose training frameworks remain predominantly Western and equaltempered. Real time accompaniment sharpens this gap: an AI partner must listen, adapt dynamically, and respect idiomatic microtonal structures. Streaming text to music models provide strong generative capabilities but lack precise control interfaces. We present MazzikaAI, a knowledge based system that uses natural language as the actuator of a realtime control loop. By compiling live MIDI, gesture, and inferred harmony into continuously updated text prompts, MazzikaAI steers an unmodified streaming generator, Google Lyria RealTime, without requiring model finetuning. The system embeds expert knowledge of six core maqamat, characteristic ornaments, and ensemble dynamics, maintaining realtime responsiveness with subsecond keytoaudibleupdate latency. Empirical evaluations demonstrate that dynamic prompt compilation reliably grounds generation in microtonal scales, significantly increasing offgrid quartertone content over baseline generation. Beyond its core implementation, MazzikaAI illustrates how deterministic knowledgebased rules can effectively bridge expert, nonWestern musical traditions and unfinetuned foundation models. This architecture establishes a scalable paradigm for realtime humanAI cocreation, offering a generalizable blueprint for interactive accompaniment, adaptive music education, and culturally inclusive generative audio across diverse global idioms.

↑