发表机构
National Yang Ming Chiao Tung University(国立阳明交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对节奏游戏图表生成,引入ChartGenEval评估框架,通过自动且经损坏测试的核心,在开放音符选择时锚定歌曲节奏,经多组测试及压力测试,能提供多维反馈,为生成器比较与迭代提供依据。
AI 中文摘要
生成的节奏游戏图表不必重现一个官方音符序列,许多音符选择都能适配同一首歌曲和难度。因此,参考音符一致性衡量的是重构,而非整个设计问题。我们引入了ChartGenEval,这是一个具有自动、经过损坏测试核心的六问题评估框架。它在将节奏锚定到歌曲的同时,让音符选择保持开放,匹配的官方图表仅提供其编写的节奏映射,而非目标音符。我们用剂量控制的故障测试每个核心输出,而非假设熟悉的统计量能衡量图表质量。在80个保留歌曲组中,七个输出轴在九次非冗余测试中满足预先指定的灵敏度和不变性标准。对40首歌曲开发面板的补充压力测试揭示了两个更广泛的教训。全图表相位估计能恢复15、30和60毫秒的注入偏移,而仅图表输出基本不变。通用模式重写使平均语言模型困惑度降低37%,循环折叠使平均自相似性提高62%。ChartGenEval因此报告单独的、特定角色的信号,而非一个代理或总分。此配置文件为比较和迭代生成器提供自动反馈;在特定任务压力测试后,选定的输出是候选优化目标或约束。
英文摘要
A generated rhythm-game chart need not reproduce one official note sequence: many note choices can fit the same song and difficulty. Reference-note agreement therefore measures reconstruction, not the full design problem. We introduce ChartGenEval, a six-question evaluation framework with an automatic, corruption-tested core. It leaves note choice open while anchoring timing to the song: the matched official chart supplies only its authored timing map, never target notes. We test each core output with dose-controlled failures rather than assume that a familiar statistic measures chart quality. Across 80 held-out song groups, seven output axes satisfy prespecified sensitivity and invariance criteria in nine nonredundant tests. Complementary stress tests on the 40-song development panel expose two broader lessons. A chart-wide phase estimate recovers injected shifts of 15, 30, and 60 ms while chart-only outputs remain essentially unchanged. Common-pattern rewriting lowers mean language-model perplexity by 37%, and loop collapse raises mean self-similarity by 62%. ChartGenEval therefore reports separate, role-specific signals instead of one proxy or total score. This profile provides automatic feedback for comparing and iterating generators; selected outputs are candidate optimization targets or constraints after task-specific stress testing.