发表机构
Japan Advanced Institute of Science and Technology(日本先进科学技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究音乐游戏关卡程序生成问题,受基于事件的符号音乐建模启发,提出令牌级序列公式,构建Transformer模型,在事件级评估中优于基线,能系统分析音频对节奏对齐事件预测的支持。
AI 中文摘要
音乐游戏关卡的程序生成是一个令人兴奋但具有挑战性的问题,因为关卡必须将音乐结构转化为定时游戏玩法事件的交互序列。大多数现有方法通过基于帧的表示来制定此任务,将音频划分为均匀的时间网格并预测每一帧的事件。这使得游戏玩法事件在许多帧中是隐含的。因此,很难描述在人工制作的关卡中发现的事件级定时关系和更长范围的结构。我们使用程序生成作为实际设置来研究音乐线索如何映射到交互事件序列。受基于事件的符号音乐建模的启发,我们提出了一种令牌级序列公式,将关卡生成视为多模态序列到序列问题。基于音频片段和关卡元数据,模型生成交替出现游戏玩法事件和节拍转移令牌的令牌序列。这在节拍空间中明确表示了动作及其相对时间。基于此公式,我们构建了一个Transformer模型。在事件级评估下,它优于代表性的基于帧的基线。它还能够系统地分析音频如何支持超越元数据条件的节奏对齐事件预测。
英文摘要
Procedural generation of music game levels is an exciting yet challenging problem, as levels must translate musical structure into interactive sequences of timed gameplay events. Most existing approaches formulate this task by frame-based representations, dividing audio into uniform time grids and predicting events at each frame. This makes gameplay events implicit across many frames. As a result, it is hard to describe event-level timing relations and longer-range structure found in human-authored levels. We use procedural generation as a practical setting to study how musical cues map to interactive event sequences. Inspired by event-based symbolic music modeling, we propose a token-level sequence formulation that casts level generation as a multimodal sequence-to-sequence problem. Conditioned on an audio excerpt and level metadata, the model generates a token sequence alternating gameplay-event and beat-shift tokens. This explicitly represents actions and their relative timing in beat space. Based on this formulation, we build a Transformer model. It outperforms representative frame-level baselines under event-level evaluation. It also enables systematic analysis of how audio supports rhythm-aligned event prediction beyond metadata conditioning.
CommentsCamera-ready version, published at ICMR 2026
Journal refProceedings of the International Conference on Multimedia Retrieval (ICMR '26), June 16-19, 2026, Amsterdam, Netherlands