连续交互扩散:面向异步工具增强推理的原生扩散运行时
Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning
浏览论文内容
中文总结 AI 辅助
针对扩散语言模型异步工具增强推理的局限,提出连续交互扩散架构,将工具交互整合至迭代去噪,可提升效率并减少冗余,首次聚焦只读工具。
中文摘要 AI 辅助
大型语言模型越来越依赖外部工具获取最新信息、执行计算并与外界交互。对于自回归模型而言,工具使用自然契合其生成过程:模型输出工具调用、等待结果后继续生成。然而,扩散语言模型(dLLMs)通过并行细化输出的多个部分进行推理,这种“停止-恢复”交互模式会造成不必要的限制,可能迫使模型在推理稳定前做出工具决策、延迟有用观测至离散调用完成、引入冗余细化与工具执行,进而损害任务准确率和推理效率。我们提出连续交互扩散(CID),一种将工具交互整合至迭代去噪过程的原生扩散模型-运行时架构。CID 分离出模型只读事实通道、由类型化认知张量表示的思维通道及显示通道。信息需求可在文本或 JSON 调用完全序列化前产生,使感知绑定能在去噪持续时启动外部读取;返回结果会被投影至演化的思维状态,可修正早期认知与显示区域。持久绑定可复用静态结果而无需重复外部执行,并按需刷新变化源。CID 旨在更早暴露证据、将工具延迟与模型计算重叠、减少重复外部工作、在新证据到达后保留有用计算。我们对该架构、运行时及训练目标进行形式化定义,并制定任务质量与端到端效率的评估协议。本论文首次聚焦只读工具,未提出实证性能主张。
英文摘要
Large language models increasingly rely on external tools to access up-to-date information, perform computation, and interact with the outside world. For autoregressive models, tool use naturally fits the generation process: the model emits a tool call, waits for the result, and then continues generating. Diffusion language models (dLLMs), however, reason by repeatedly refining many parts of their output in parallel, making this stop-and-resume interaction pattern unnecessarily restrictive. It can force tool decisions before the model's reasoning has stabilized, delay useful observations until a discrete call finishes, and introduce redundant refinement and tool execution, potentially hurting both task accuracy and inference efficiency. We introduce Continuous Interaction Diffusion (CID), a diffusion-native model--runtime architecture that integrates tool interaction into iterative denoising. CID separates a model-read-only fact channel, a thought channel represented by a Typed Cognitive Tensor, and a display channel. Information needs can emerge before a textual or JSON call is fully serialized, allowing perceptual bindings to launch external reads while denoising continues. Returned results are projected into the evolving thought state and can revise earlier cognition and display regions. Persistent bindings reuse static results without repeated external execution and refresh changing sources when needed. CID is designed to expose evidence earlier, overlap tool latency with model computation, reduce duplicate external work, and preserve useful computation after new evidence arrives. We formalize the architecture, runtime, and training objectives, and define an evaluation protocol for task quality and end-to-end efficiency. This first paper focuses on read-only tools and makes no empirical performance claims.
发表机构
- Nanjing University(南京大学)
机构由 AI 辅助整理,请以论文原文为准。