arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11299cs.AIcs.SD

DuplexAgent-RSI:用于全双工语音智能体协作的递归工具改进

DuplexAgent-RSI: Recursive Harness Improvement for Full-Duplex Voice Agent Collaboration

Yingda Shen, Yuxiang Wang, Kunyu Feng, Qinke Ni, Jiaqi Li, Minghao Hsu, Junan Zhang, Dekun Chen, Yutong Bian, Zhizheng Wu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出DuplexAgent全双工语音智能体协作系统及Duplex-Harness-RSI递归工具改进闭环,通过模块化设计结合全双工交互与复杂任务执行,在基准测试中表现优于对比系统。

中文摘要 AI 辅助

语音智能体正趋向一种协作模式:全双工交互模型作为对话入口驻留在实时通道中,而搜索、推理和编码则通过异步委托处理。全双工模型支持连续的听与说,但复杂推理和工具使用可能超出其能力范围;编码智能体可规划并执行扩展任务,但其顺序接口不适合实时对话。将二者结合需要一个工具,该工具需协调任务接受、进度、取消、替换及结果交付,同时保持对话响应性。现有工具常依赖耦合启发式方法,难以从证据中系统改进。我们提出DuplexAgent,这是一个全双工协作系统,其工具将工作流表示为六个可编辑模块;还提出Duplex-Harness-RSI,这是一个从交互轨迹中修改这些模块的闭环。模拟器自动生成定时测试对话,运行系统并生成故障轨迹,以识别需要修复的协作模块。委托池中的推理大语言模型和编码智能体也服务于改进循环:测试规划器从观察到的弱点和修复档案中选择下一个测试,工具编辑器提出针对性的模块更改。服务用户的能力也可改进系统的协作。在智能、智能体和全双工基准上的实验表明,DuplexAgent将连续交互与困难推理及复杂任务执行相结合,比对比的委托系统实现了更强的口语知识和可执行工具分数,同时保持了强大的中断响应。工具消融进一步表明,这种模块化、可验证的循环优于初始工具和缺乏诊断及修复档案的重复编辑。

英文摘要

Voice agents are converging on a collaboration pattern: a full-duplex interaction model stays on the live channel as the entry to the conversation, while search, reasoning, and coding are handled through asynchronous delegation. A duplex model supports continuous listening and speaking, but complex reasoning and tool use may exceed its capabilities. A coding agent can plan and execute extended tasks, but its sequential interface is a poor fit for live conversation. Combining them requires a harness that coordinates task acceptance, progress, cancellation, replacement, and result delivery while keeping the conversation responsive. Existing harnesses often rely on coupled heuristics, making them difficult to improve systematically from evidence. We present DuplexAgent, a full-duplex collaboration system whose harness expresses this workflow as six editable modules, and Duplex-Harness-RSI, a closed loop that revises them from interaction traces. A simulator automatically generates timed test conversations, runs the system, and produces failure traces that identify the collaboration modules requiring repair. Reasoning LLMs and coding agents in the delegation pool also serve the improvement loop: the Exam Planner selects the next tests from observed weaknesses and the repair archive, and the Harness Editor proposes targeted module changes. The capabilities that serve the user thus also improve the system's coordination. Experiments on intelligence, agentic, and duplex benchmarks show that DuplexAgent combines continuous interaction with difficult reasoning and complex task execution, achieving stronger spoken-knowledge and executable-tool scores than the compared delegated systems while maintaining strong interruption response. A harness ablation further shows that this modular, verifiable loop outperforms the initial harness and repeated editing that lacks its diagnosis and repair archive.

发表机构

  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Amphion Technology Co., Ltd(Amphion科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑