arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31937cs.SEcs.AI

验证作为LLM智能体的架构层:V模型设计及其确定性核心的初步研究

Verification as an Architectural Layer for LLM Agents: A V-Model Design, and a Pilot Study of Its Deterministic Core

Ali Afoud, Jie JW Wu

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM智能体生成循环无法拒绝输出的问题,提出将验证作为架构层,采用V模型设计,以确定性核心实现故障定位与主动停止,初步实验表明验证配置显著提升任务完成率。

中文摘要 AI 辅助

基于ReAct模式构建的大语言模型(LLM)智能体将四项职责集中于一个模型中:选择策略、选择每个动作、格式化动作以及判断结果是否充分。生成循环之外没有任何机制可以拒绝其输出,因此无法取得进展的智能体不会报告失败;它会一直运行直到外部预算将其终止。我们提出将验证作为架构层,借鉴软件工程中的V模型:规格层级从需求下降到各个步骤,每个层级配有一个专用验证器,一个确定性控制器强制执行每个判定,并且只有验证结果写入记忆,因此拒绝可以定位到引入故障的层级,智能体通过拒绝而非耗尽资源来停止。每个验证器将零成本的确定性“门”与可选的LLM“裁判”分开,因此可以独立衡量各自的贡献和成本。我们报告了一项初步研究,实现了验收级和单元级验证器对,比较了五种配置,这些配置共享一个执行器、工具集和评分器,仅在验证上有所不同,使用8B参数骨干模型在MuSiQue的四跳层上进行。在47次执行中,两种未经验证的配置未能回答十个问题中的任何一个,每次运行都终止于步骤上限或提供者令牌限制;没有规划器的验证配置回答了八个问题,并在其余问题上弃权(不执行)。确定性门在零边际成本下产生了九次观察到的修正中的八次,而一旦存在验证,规划反而降低了性能。这些结果表征了终止行为,而非大规模准确性;我们概述了一个十二个月的计划,以完成并评估完整架构,包括初步研究省略的集成级对。

英文摘要

Large language model (LLM) agents built on the ReAct pattern concentrate four responsibilities in one model: selecting a strategy, choosing each action, formatting it, and judging whether the result is adequate. Nothing outside the generative loop can reject its output, so an agent that cannot make progress does not report failure; it runs until an external budget stops it. We propose treating verification as an architectural layer by adapting the V-model from software engineering: specification levels descend from requirements to individual steps, each level is paired with a dedicated verifier, a deterministic controller enforces every verdict, and only verification outcomes write to memory, so a rejection localizes the level that introduced the fault and an agent halts by declining rather than by exhaustion. Each verifier separates a zero-cost deterministic \emph{gate} from an optional LLM \emph{judge}, so the contribution and cost of each can be measured independently. We report a pilot implementing the acceptance- and unit-level verifier pairs, comparing five configurations that share one executor, tool set, and scorer and differ only in verification, on the four-hop stratum of MuSiQue with an 8B-parameter backbone. Across 47 executions, the two unverified configurations answered none of ten questions, every run ending at a step cap or provider token limit; the verified configuration without a planner answered eight and abstained on the rest. Deterministic gates produced eight of the nine observed corrections at zero marginal cost, and planning degraded performance once verification was present. These results characterize termination behavior, not accuracy at scale; we outline a twelve-month plan to complete and evaluate the full architecture, including the integration-level pair the pilot omits.

发表机构

  • Michigan Technological University(密歇根理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑