arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30662cs.AI

LLM帕金森综合征:执行控制失败、令牌低效持续行为,以及面向自主语言模型代理的不确定性感知全局执行控制架构

LLM Parkinsonism: Executive-Control Failure, Token-Inefficient Persistence, and an Uncertainty-Aware Global Executive Control Architecture for Autonomous Language-Model Agents

Dongsheng Xiao, Zeyuan Wang, Xuzhe Xia, Bo Zhao, Yankai Cao

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM代理在目标达成后仍持续行动的问题,提出不确定性感知的全局执行控制架构GEC v0.2,分离行动生成与项目级控制,在保持成功率的同时显著降低令牌使用和复杂性。

中文摘要 AI 辅助

大型语言模型(LLM)能够进行规划、使用工具、编写代码并执行长期工作流程,然而,强大的局部能力并不能保证项目级别的执行控制。代理可能在原始目标已满足后继续行动,产生低价值的改进、重复验证,以及修复自身造成的复杂性。我们使用LLM帕金森综合征作为一个狭义定义的非临床隐喻,来描述这种尽管任务级价值递减仍持续行动的模式。我们认为,这一问题不能仅由自回归下一个令牌预测来解释,而更直接地归因于将提案生成、范围解释、进度评估和停止权限集中在同一个自条件循环中。因此,我们引入了全局执行控制(GEC)v0.2,一种不确定性感知的治理架构,将行动生成与项目级控制分离。在一个包含24,000个回合的匹配候选基准测试中,在常见的40,000令牌上限下,首个候选基线实现了67.42%的硬目标成功率,候选集局部控制实现了96.53%,而GEC实现了96.57%。候选集控制表明,访问多个候选行动解释了大部分成功增益;相对于该控制,GEC在保持成功率的同时,将平均令牌使用量从19,782降至12,574(减少36.4%),并将40,000令牌上限下完成时的平均令牌数从16,136限制至13,114(减少18.7%),同时消除了测量到的完成前漂移并大幅降低了总体复杂性。治理开销敏感性在每周期额外500个合成治理令牌的情况下仍保持有利。这些机制模拟支持对范围、证据、资源使用和停止的显式治理,而实时模型验证仍然是必要的。

英文摘要

Large language models (LLMs) can plan, use tools, write code, and execute long-horizon workflows, yet strong local competence does not guarantee project-level executive control. Agents may continue acting after the original objective is satisfied, producing low-value refinements, repeated verification, and repairs to self-created complexity. We use LLM Parkinsonism as a narrowly defined, non-clinical metaphor for this pattern of persistent action despite diminishing task-level value. We argue that the problem is not explained by autoregressive next-token prediction alone, but more directly by concentrating proposal generation, scope interpretation, progress assessment, and stopping authority within the same self-conditioned loop. We therefore introduce Global Executive Control (GEC) v0.2, an uncertainty-aware governance architecture that separates action generation from project-level control. In a 24,000-episode matched-candidate benchmark under a common 40,000-token ceiling, a first-candidate baseline achieved 67.42% hard-goal success, a candidate-set local control achieved 96.53%, and GEC achieved 96.57%. The candidate-set control shows that access to multiple candidate actions explains most of the success gain; relative to that control, GEC preserved success while reducing mean token use from 19,782 to 12,574 (36.4%) and restricted mean tokens to completion at the 40,000-token ceiling from 16,136 to 13,114 (18.7%), while eliminating measured pre-completion drift and sharply reducing gross complexity. Governance-overhead sensitivity remained favorable through an additional 500 synthetic governance tokens per cycle. These mechanistic simulations support explicit governance of scope, evidence, resource use, and stopping, while live-model validation remains necessary.

发表机构

  • Queensland Brain Institute, The University of Queensland(昆士兰大学昆士兰脑研究所)
  • Northwestern Polytechnical University(西北工业大学)
  • The University of British Columbia(不列颠哥伦比亚大学)
  • Neurointelligence Labs(神经智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑