arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17416cs.AI

永不停思:连续时间语言智能体

Never Stop Thinking: Continuous-Time Language Agents

  • Pine AI
  • University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

Bojie Li, Noah Shi

AI总结:

本文提出连续时间语言智能体,通过中断-恢复编排实现边听边思、边说边思,降低延迟,并引入ReactiveBench基准,证明可验证信号与在线RL优化使连续思考有益于任务完成。

AI中文摘要:

基于LLM的语音智能体遵循僵化的听-思-说循环,在每次回复前插入数秒的静默延迟。我们证明,在轻量级中断-恢复编排器下,连续时间认知(边听边思和边说边思)可从未经修改的文本模型中涌现,将实时流水线延迟整体降低19%,在机制针对的场景中降低一半。为衡量连续时间思考是否改善智能体完成任务的能力,我们引入ReactiveBench:120个交互场景,依据预注册的二元需求评分,外加一个可验证的流式轨道,依据精确正确性评分。ReactiveBench揭示了一个影响广泛的陷阱:LLM评判者奖励可见推理;连续时间思考在独立评判者下其大规模评判的“优势”符号反转,且经评判训练过的模型在思考时客观完成的需求更少。随后,一项五阶段训练研究在三个层面定位了正确信号。其来源:可验证目标将思考从有害转为有益。其结构:统一奖励遗漏的任何内容,优化都会将其权衡掉;处处简洁会侵蚀多跳工具链。其优化器:偏好优化只能权衡冲突的子目标,而基于类型形状奖励的在线RL同时改善每个正确性轴,将流式完成率从48%提升至跨种子的73±5%,并在更大规模及第二个模型上得到复现。编排使连续时间交互成为可能;一个正确来源、塑造和优化的可验证信号,使其变得优秀。

英文摘要:

Voice agents built on LLMs follow a rigid listen-think-speak loop that inserts seconds of dead air before every reply. We show that continuous-time cognition (thinking while listening and thinking while speaking) emerges from an unmodified text model under a lightweight interrupt-and-resume orchestrator, cutting live-pipeline latency by 19% overall and by half in the regime the mechanism targets. To measure whether continuous-time thinking improves what agents accomplish, we introduce ReactiveBench: 120 interactive scenarios scored against pre-registered binary requirements, plus a verifiable streaming track scored by exact correctness. ReactiveBench exposes a pitfall with broad consequences: LLM judges reward visible reasoning; a large judged "advantage" of continuous-time thinking reverses sign under an independent judge, and judge-trained models objectively complete fewer requirements when they think. A five-stage training study then locates the right signal at three levels. Its source: verifiable objectives turn thinking from harmful to helpful. Its structure: whatever a uniform reward omits, optimization trades away; brevity everywhere erodes multi-hop tool chaining. Its optimizer: preference optimization can only trade conflicting sub-goals against each other, while on-policy RL over a type-shaped reward improves every correctness axis at once, raising streaming completion from 48% to 73+/-5% across seeds and replicating at larger scale and on a second model. Orchestration makes continuous-time interaction possible; a verifiable signal, correctly sourced, shaped, and optimized, makes it good.

↑