发表机构
Xiamen University; Fuyao University of Science and Technology(厦门大学; 福耀科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对真实场景音频增强的复杂失真耦合与个性化需求难题,提出基于MLLM的智能体StrixAE,经两阶段训练后性能优于多数现有方案,实现了更优的音频增强效果与鲁棒性。
AI 中文摘要
真实场景中的音频增强涉及复杂的失真耦合,且需要个性化增强处理,现有解决方案难以同时满足这两个需求。为提升在这类场景下的鲁棒性并实现自主运行,我们提出StrixAE,一种基于多模态大语言模型(MLLM)的智能体。StrixAE将MLLM作为控制器,协调多个音频增强与个性化模型。为进一步提升系统鲁棒性、减少伪影并增强在不同真实场景中的泛化能力,StrixAE通过两阶段流程训练:第一阶段是在AcoustBench上进行思维链(CoT)监督微调,以建立基础推理与工具调用能力;第二阶段是音频感知强化学习(APRL),这是一种专为音频恢复流程设计的奖励机制,可联合优化格式有效性、结构连贯性与感知质量。与通用强化学习微调不同,APRL引入结构化奖励以确保流程可执行与逻辑环节有序,使智能体能生成可靠、可解释的增强方案,且不会出现工具幻觉。基于真实世界测试数据集,我们提出的方法优于大多数现有开源及专有解决方案,在多项感知指标上达到了最先进性能,并展现出强大的泛化鲁棒性。
英文摘要
Audio enhancement in real-world scenarios involves complex distortion couplings and requires personalized enhancement. Existing solutions struggle to address both simultaneously. To improve robustness and enable autonomous operation in such scenarios, we propose StrixAE, an agent based on a multimodal large language model (MLLM). StrixAE leverages the MLLM as a controller to coordinate multiple audio enhancement and personalization models. To further enhance system robustness, reduce artifacts, and improve generalization across diverse real-world scenarios, StrixAE is trained through a two-stage process: first, CoT supervised fine-tuning on AcoustBench to ground basic reasoning and tool invocation; second, Audio Perception Reinforcement Learning (APRL), a reward design specifically tailored for audio restoration pipelines that jointly optimizes format validity, structural coherence, and perceptual quality. Unlike generic RL fine-tuning, APRL introduces structured rewards that enforce executable pipelines and logical section ordering, enabling the agent to produce reliable, interpretable enhancement plans without hallucinated tools. Based on real-world test datasets, our proposed method outperforms most existing open-source and proprietary solutions, achieving state-of-the-art performance across multiple perceptual metrics and demonstrating strong generalization robustness.