发表机构
Technical University of Darmstadt; UCL Centre for AI(达姆施塔特工业大学; 伦敦大学学院人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对扩散语言模型智能体重试循环问题,提出任务盲目破坏模型与不变性命题,并实例化Reflect Reverse反向评分方法,在四个多轮具身基准上提升任务成功率和进展率。
AI 中文摘要
基于扩散的大语言模型(dLLMs)承诺通过并行解码打破自回归智能体的顺序延迟瓶颈,但最近的评估表明这种效率并未转化为具身智能体的能力:dLLM支持的智能体反复陷入重试循环,在动作失败很久后仍重新发出该动作。我们给出了这一失败的机制解释和无需训练的治疗方法。我们将重试循环追溯到掩码解码的自适应性:采样器提交它最有信心的位置,并推迟不确定的位置,而在失败状态下,上下文已经为被推迟的决策提供了有信心的填充,即失败的动作本身,因此重试被提交,而失败反馈从未被面对。我们将由此产生的动作分布扭曲建模为任务盲目的破坏:上下文显著的动作(例如,刚采取的动作)获得膨胀的概率,其因子取决于状态和动作,但不取决于任务。在此模型下,我们分析了一个不变性命题:任务盲目因子从反向条件中精确抵消,即给定状态和候选动作的任务似然,这与理想未破坏模型的任务后验一致。掩码dLLMs通过掩码任务标记和去噪来原生评估反向条件,不同于自回归模型,代价是每个候选需要几次并行传递。我们将该规则实例化为Reflect Reverse,并在四个多轮具身基准上评估,它在任务成功率和进展率上优于前向评分基线。
英文摘要
Diffusion-based large language models (dLLMs) promise to break the sequential latency bottleneck of autoregressive agents through parallel decoding, but recent evaluations show this efficiency does not transfer to embodied agentic competence: dLLM-backed agents repeatedly fall into retry loops, re-issuing an action long after it has failed. We give a mechanistic account of this failure and a training-free remedy. We trace the retry loop to the adaptivity of masked decoding: the sampler commits the positions it is most confident about and defers the uncertain ones, and at a failure state the context already offers a confident fill for the deferred decision, i.e. the failed action itself, so the retry is committed without the failure feedback ever being confronted. We model the resulting distortion of the action distribution as a task-blind corruption: contextually salient actions (e.g., the action just taken) receive inflated probability by a factor that depends on the state and the action but not on the task. Under this model, we analyse an invariance proposition: the task-blind factor cancels exactly from the reverse conditional, i.e. the likelihood of the task given the state and a candidate action, which coincides with the task posterior of an idealized uncorrupted model. Masked dLLMs evaluate the reverse conditional natively, unlike autoregressive models, by masking the task tokens and denoising, at the cost of a few parallel passes per candidate. We instantiate the rule as Reflect Reverse and evaluate it on four multi-turn embodied benchmarks, where it improves task success and progression rates over forward-scoring baselines.