任务状态应在多大程度上影响 LLM 智能体?
How Strongly Should Task State Influence an LLM Agent?
浏览论文内容
中文总结 AI 辅助
本研究系统比较了任务状态以文本展示、指令告知和门控强制执行三种方式对LLM智能体长时程任务可靠性的影响,发现强制执行在状态可判定且失败频繁时最有效,但受限于状态正确性,且在不同领域效果各异。
中文摘要 AI 辅助
长时程分配的任务要求 LLM 智能体跟踪任务状态:哪些步骤已完成、被阻塞、被取消或可重复执行。智能体系统要么将状态作为提示中的文本保留并依赖模型读取该文本,要么将状态移入一个强制执行的模块中,而每个系统都是作为一个整体进行评估的,因此没有人知道可靠性在多大程度上来自状态的展示、告知或强制执行。我们固定任务规则、模型和成对回合,并改变任务状态到达智能体的强度:原始转录、精确检查清单、由简报编译且仅通过执行收据推进的状态机给出的每回合指令,或该状态机上拒绝违反状态动作的强制执行门控;每个回合都通过精确负载匹配动态真实值进行评分。在三个模型、两种推理模式和两个领域中,四个发现无需每回合推理即可成立:展示准确状态不可靠,智能体自行编写的未经验证的账本优于展示给它的准确检查清单,指令的帮助程度与模型的服从度成正比,强制执行不需要服从但受限于其状态的正确性以及将请求映射到步骤的匹配器;235B 智能体的每回合推理压缩了这些差异,但未能修复文本层面的不足。由 $\ au^2$-bench 的航空公司策略编译的同一门控,将 235B 智能体的 pass$^1$ 从 0.39 提升至 0.54,而对很少违反策略的 35B 智能体则无任何改变;在 PM-Bench 上,由于行动回合依赖于识别线索而非状态,展示记录是最佳层级——匹配或超越两个门控并逆转账本优于检查清单的发现——而强制执行匹配器的判断会使 35B 智能体低于其原始转录。当失败是状态可判定且频繁时,强制执行有效;当门控的判断错误时,强制执行有害。
英文摘要
Long-horizon assigned work requires an LLM agent to track the state of a task: which steps are done, blocked, cancelled, or open to repetition. Agent systems either keep this state as text in the prompt and rely on the model to read that text, or move the state into a module that enforces it, and each system is evaluated as a whole, so no one knows how much reliability comes from the state being shown, told, or enforced. We fix the task rules, the model, and paired episodes and vary how strongly task state reaches the agent: a raw transcript, an exact checklist, per-turn directives from a state machine compiled from the brief and advanced only by execution receipts, or an enforcement gate on that machine that refuses state-violating actions; every episode is scored by exact payload matching against dynamic ground truth. Across three models, two reasoning regimes, and two domains, four findings hold without per-turn reasoning: displaying accurate state is unreliable, an unverified ledger the agent writes itself beats an accurate checklist it is shown, directives help in proportion to the model's obedience, and enforcement needs no obedience but is bounded by the correctness of its state and by the matcher that maps requests to steps; per-turn reasoning at a 235B agent compresses these separations without repairing the text rungs. The same gate, compiled from $τ^2$-bench's airline policy, raises a 235B agent's pass$^1$ from 0.39 to 0.54 and changes nothing for a 35B agent that rarely violates the policy; on PM-Bench, where acting turns on recognizing a cue rather than on state, showing the record is the best rung--matching or beating both gates and reversing the ledger-over-checklist finding--and enforcing the matcher's judgement drops a 35B agent below its raw transcript. Enforcement pays when failures are state-decidable and frequent, and hurts when the gate's judgement is wrong.
发表机构
- University of Waterloo(滑铁卢大学)
- Sungkyunkwan University(成均馆大学)
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。