发表机构
Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于事件级变换提示的音效补全方法,利用预训练模型微调修复缺失音效,在保留皮肤上优于通用音频编辑器。
AI 中文摘要
为新的游戏角色皮肤创建音效需要独特的声学标识,同时保留游戏事件的角色。挑战在于完成一组连贯的相关音效,而这些音效所需的重新设计程度各不相同。我们将此任务表述为基于基础皮肤音频、已完成的目标资产和文本设计描述的条件补全。我们开发了一个流程来收集、处理和对齐《英雄联盟》皮肤中的对应事件。基于Stable Audio 3的预训练音频先验,我们微调了一个潜在修复模型以联合补全缺失事件。一个有符号的软保留掩码编码了可用音频,并为每个缺失事件提供可调整的变换提示,指定保留与重新设计之间的所需平衡。在保留皮肤上的实验表明,与评估的通用音频编辑器相比,重建效果有所改进。目标派生的提示进一步提高了配对相似性,三级提示保留了连续指导的大部分益处。
英文摘要
Creating sound effects for a new game-character skin requires a distinct acoustic identity while preserving gameplay-event roles. The challenge is to complete a coherent set of related sounds whose required degrees of redesign differ. We formulate this task as completion conditioned on base-skin audio, completed target assets, and a textual design description. We develop a pipeline to collect, process, and align corresponding events across League of Legends skins. Building on Stable Audio 3's pretrained audio prior, we fine-tune a latent inpainting model to jointly complete missing events. A signed soft retention mask encodes available audio and an adjustable transformation hint for each missing event, specifying the requested balance between retention and redesign. Experiments on held-out skins show improved reconstruction over the evaluated general-purpose audio editors. Target-derived hints further improve paired similarity, with three-level hints retaining most of the benefit of continuous guidance.
Comments5 pages, 1 figure, 3 tables. Submitted to ICASSP 2027