发表机构
Shanda Group(盛大集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过AgenticTTS-Forge工作流验证LLM代理可自动化TTS强化学习配方,恢复并改进配方,但暴露数据泄露、代理膨胀等测量陷阱,表明核心约束在于测量而非推理。
AI 中文摘要
尽管强化学习(RL)后训练修复了零样本文本到语音(TTS)的局部段错误,但获得有效的配方仍依赖繁琐的手动调整,且LLM代理能否接管这一研究流程尚不清楚。我们通过AgenticTTS-Forge(一种围绕共享工作区构建人类指导和代理执行的协作工作流)对这一问题进行了研究,并将其应用于CosyVoice2-0.5B。为衡量代理自动化的内容,我们逐阶段对照已发布配方审计其轨迹;为衡量其利用的内容,我们使用对代理隐藏的保留观察者对其策略进行评分。结果表明,代理恢复了一个未明确指定的配方,对其进行了改进,并在收益停滞时,未经提示地调研文献,从LM载体转向流载体,将Bad案例减半。然而,其自主性暴露了数据、代理和算法三个轴上的陷阱:保留集通过合同从未读取的渠道泄露,自塑奖励在评分处膨胀代理指标,且单独调优的策略无法加性组合。这些发现表明,约束在于测量而非推理,并可为设计合同读取代理所读每个渠道的框架提供参考。
英文摘要
Although reinforcement learning (RL) post-training repairs the localized segmental errors of zero-shot text-to-speech (TTS), arriving at a working recipe still relies on tedious manual tuning, and whether LLM agents can take over this research pipeline is unclear. We investigate this question with AgenticTTS-Forge, a collaborative workflow that structures human guidance and agentic execution around a shared workspace, applied to CosyVoice2-0.5B. To measure what the agent automates, we audit its trajectory stage by stage against the published recipe. To measure what it exploits, we score its policies with held-out observers hidden from the agent. Our results show that the agent recovers an underspecified recipe, improves it, and, when gains stall, surveys the literature unprompted and pivots from the LM carrier to the flow carrier, halving Bad cases. However, its autonomy exposes three traps across the data, proxy, and algorithm axes: the held-out set leaks through a channel the contract never reads, a self-shaped reward inflates the proxy where it is scored, and separately tuned policies do not compose additively. These findings show that the binding constraint is measurement rather than reasoning, and can inform the design of harnesses whose contracts read every channel the agent does.
CommentsPreprint, work in progress