arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SpeechGym:一种用于通过强化学习训练语音智能体的原生音频 Gym

SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning

Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar, Jia-Hong Huang, Qi Luo, M. Maruf, Ivan Bulyko, Ge Liu, Roger Ren

arXiv 2608.26432首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Amazon AGI Foundations(伊利诺伊大学厄巴纳-香槟分校; 亚马逊AGI基础研究部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对语音智能体训练中梯度无法流动、难以强化学习的问题,提出原生音频环境 SpeechGym,用每轮过程奖励解决稀疏性问题,使开放权重模型在语音基准上任务成功率翻倍且排名提升。

AI 中文摘要

语音智能体必须完全通过语音调用工具并进行多轮对话,然而主流范式却在文本环境中对其进行训练。现有框架要么在专有语音 API 周围级联 TTS 和 ASR,其中梯度无法流动且每次调用的成本使得基于策略的强化学习难以实现;要么停留在文本环境中:它们可以评估语音智能体,但无法对其进行改进。我们提出了 SpeechGym,这是一种原生音频智能体环境,其中两个全模态模型以原生音频形式进行对话,无需外部 ASR 或 TTS,也无 API 边界,基于已建立的文本智能体基准的未修改任务、工具和成功检查,因此交互模态是唯一变量,且循环保持局部性并可端到端训练。音频智能体能力并非来自音频理解。语音引入的失败是感知而非推理缺陷:智能体选择了正确的工具和正确的参数槽,但用波形中听错的值填充它,而这单一错误会级联为调用失败、同一调用的重试以及浪费的步骤预算。第二个失败是行为性的:在坚持呼叫者的要求下,智能体执行了未授权的写入,并在结束回合时认为自己提供了帮助。这两种失败都可训练,因为环境会免费对其进行标记:参数听错的调用会在数据库中失败,而参数正确的调用则会成功。障碍是稀疏性,而非信号。仅基于结果的 GRPO 在此处梯度匮乏,因为几乎每个 rollout 组都以相同方式失败,而每轮过程奖励(对每个成功的工具调用进行奖励)则为几乎每个组恢复了方差。通过这种方式训练后,智能体无需进一步微调即可迁移到独立实现的语音基准,使任务成功率提高了一倍以上,并将一个开放权重模型从该排行榜的最后一名提升至第二名,同时使用的轮次和 token 比训练前更少。

英文摘要

Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing frameworks either cascade TTS and ASR around a proprietary voice API, where gradients cannot flow and per-call cost makes on-policy reinforcement learning prohibitive, or stay in text: they measure voice agents but cannot improve them. We present SpeechGym, an audio-native agentic environment in which two omni-modal models converse in native audio, with no external ASR or TTS and no API boundary, over the unmodified tasks, tools and success check of an established text agentic benchmark, so that the interaction modality is the only variable and the loop stays local and trainable end to end. Audio agentic capability does not follow from audio understanding. The failures speech introduces are perceptual rather than reasoning deficits: the agent picks the right tool and the right argument slot but fills it with a value misheard from the waveform, and that single error cascades into a failed call, a retry of the same call, and a wasted step budget. A second failure is behavioural: under an insistent caller the agent performs an unauthorised write and ends the episode believing it helped. Both are trainable, because the environment labels them for free: a call with a misheard argument fails against the database while a correct one succeeds. The obstacle is sparsity, not signal. Outcome-only GRPO is gradient-starved here, since almost every rollout group fails identically, while a per-turn process reward crediting each successful tool call restores variance to nearly every group. Trained this way, the agent transfers with no further tuning to an independently implemented voice benchmark, more than doubling task success and carrying an open-weights model from last place to second on that leaderboard, while using fewer turns and tokens than before training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑