谁执笔?让规范而非智能体来签字
Who Holds the Pen? Let Specifications, Not Agents, Sign Off
- The University of Texas at Arlington(德克萨斯大学阿灵顿分校)
- Monash University(莫纳什大学)
- Kent State University(肯特州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对大语言模型智能体在无独立规范权威下执行与完成声明不可靠的问题,提出SpecHarness框架,将规范编译为义务状态以治理执行,实验显示其能提升合规性。
AI中文摘要:
大语言模型智能体日益在单一智能体循环中结合生成、决策、执行和自我评估。尽管它们在外部规范(如任务指令、指南、输出模式和可复用技能)下运行,这些规范通常仍是同一模型(该模型既执行又声明完成)的上下文,从而不存在独立的规范权威边界。我们识别出由此产生的两个缺口。理解-执行缺口出现在需求被理解但未在执行中得到满足时;状态-权威缺口出现在智能体的解释或完成声明未确立所需状态时。在SkillsBench上,仅使用智能体可见的提示、工作区信息和注入的技能规范,我们提取了509条源自源头的任务指令。在七个模型上,仅有79.6%–86.4%被满足,而完成声明率超过官方评估器通过率28.7–37.9个百分点。因此,我们将智能体提案与权威状态分离。智能体可以规划、行动和请求完成,但只有来自合格提供者的可接受证据才能确立规范治理的状态。SpecHarness通过将可见规范编译为源链接的义务,并通过版本化的义务状态治理执行和最终化,来操作化这一原则。可验证的需求在运行时被调解或验证,而模糊或主观的需求仍保持咨询性质。在遵循指南和工件生成任务上的实验表明,规范不仅可以作为行为指导,还可以作为合规执行和完成的权威。
英文摘要:
Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the same model that acts and declares completion, leaving no independent specification authority boundary. We identify two resulting gaps. The understanding--execution gap arises when a requirement is understood but not satisfied in execution; the state--authority gap arises when an agent's interpretation or completion claim does not establish the required state. On SkillsBench, using only agent-visible prompts, workspace information, and injected skill specifications, we extract 509 source-grounded task directions. Across seven models, only 79.6%--86.4% are satisfied, while completion-claim rates exceed official evaluator pass rates by 28.7--37.9 percentage points. We therefore separate agent proposals from authoritative state. Agents may plan, act, and request completion, but only admissible evidence from qualified providers may establish specification-governed state. SpecHarness operationalizes this principle by compiling visible specifications into source-linked obligations and governing execution and finalization through versioned obligation state. Verifiable requirements are mediated or validated at runtime, while ambiguous or subjective requirements remain advisory. Experiments on guideline-following and artifact-generation tasks show that specifications can serve not merely as behavioral guidance, but as authority over compliant execution and completion.