AI 中文总结
本研究通过实证探究harness工程中确定性执行约束对LLM智能体的影响,发现添加结构化规划可提升可复现性与任务成功率,但延迟存在模型依赖的成本差异。
AI 中文摘要
基于大语言模型(LLM)的智能体即便在给定相同任务和工具时,也会表现出显著的运行间执行差异——这在探索性使用中可接受,但在金融和合规等受监管领域中不可接受。本研究探讨harness工程:将智能体封装在确定性执行层(有限状态控制、强制工具选择、输出验证、有限重试及结构化规划)中,测量其对执行确定性和任务成功率的影响。在两个合成任务(金融与法律)和两个开放权重模型(Qwen-2.5-7B-Instruct、Gemma-3-27B)上,初始harness产生的结果不一:它在四个模型-任务单元中的一个显著提升了可复现性,在两个单元中显著降低,在第四个单元中无影响。通过轨迹级诊断发现了原因:一旦工具序列、状态序列和输出已高度一致,无约束的自由文本规划步骤便成为剩余方差的主要来源。添加结构化规划——在调用任何工具前依据固定模式验证规划——可完全消除该影响:四个单元中的三个在N=100时达到可复现率和确定性指数1.000,且四个单元中的三个任务成功率升至100%。这种提升并非无代价:所有单元的token成本均下降,但延迟呈现出真正的、样本量稳健的模型差异——一个模型在约束下更快,另一个则明显变慢。我们认为harness工程是提升智能体可靠性的独特有效学科,但其成本依赖模型,必须进行测量而非假设。
英文摘要
Large Language Model (LLM) based agents exhibit substantial run-to-run execution variance even when given identical tasks and tools -- acceptable for exploratory use but unacceptable in regulated domains such as finance and compliance. We study harness engineering: wrapping an agent in a deterministic execution layer (finite-state control, forced tool selection, output validation, bounded retry, and structured planning) and measuring its effect on execution determinism and task success. Across two synthetic tasks (finance and legal) and two open-weight models (Qwen-2.5-7B-Instruct, Gemma-3-27B), a first-pass harness produces a mixed result: it significantly improves reproducibility in one of four model-task cells, significantly degrades it in two, and has no effect in the fourth. A trace-level diagnostic finds the cause: once tool sequence, state sequence, and output are already highly consistent, an unconstrained free-text planning step becomes the dominant remaining source of variance. Adding Structured Planning -- validating the plan against a fixed schema before any tool is invoked -- eliminates the effect entirely: three of four cells reach a Reproducibility Rate and Determinism Index of 1.000 at N=100, and task success rises to 100% in three of four cells. The gain is not free: token cost falls in every cell, but latency shows a genuine, sample-size-robust split by model -- one becomes faster under the constraint, the other markedly slower. We argue harness engineering is a distinct, effective discipline for agent reliability, but its cost is model-dependent and must be measured, not assumed.
Comments9 pages, 3 tables