用于智能体框架优化的因果改进图
Causal Improvement Graph for Agentic Harness Optimization
- Nanyang Technological University(南洋理工大学)
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出因果改进图(CIG),一种图治理的元框架框架,通过外部化改进状态并利用局部提议者操作,在多种智能体任务中发现比基线更强的框架,且对求解器和提议者选择鲁棒。
AI中文摘要:
智能体框架(Agentic Harness)是构建任务上下文并控制执行流程的运行时,从而塑造整体智能体性能。在给定固定模型和外部评估的情况下,自动化框架优化旨在通过迭代的提议-评估循环来改进该运行时,以更好地解决目标任务。现有的元框架方法主要采用以提议者为中心的发现方式,其中基于LLM的提议者整合累积的实验发现以确定后续的框架修订。这随着历史记录扩展及其底层实验逻辑变得难以辨别,将维护不断演进的改进状态的责任置于提议者身上。在本文中,我们引入了因果改进图(CIG),一种图治理的元框架框架,将不断演进的改进状态外部化到持久图中,允许先前的发现通过局部提议者操作直接治理后续的框架优化。CIG增长并链接证据、假设、干预和结果节点,以表示观察到什么、如何解释、如何测试该解释以及评估揭示了什么。它们的结构关系保留了改进状态在迭代间如何变化,允许局部提议者直接基于先前发现之间的关系构建,而不是从原始历史中恢复它们。在各种智能体任务中,CIG发现了比先前元框架基线更强的框架,并且对任务求解器和提议者的选择保持鲁棒。结构消融进一步支持了具有图治理演进的显式改进状态的设计。
英文摘要:
Agentic Harness is the runtime that constructs task context and controls execution flow, thereby shaping overall agent performance. Given a fixed model and external evaluation, automated Harness optimization seeks to improve this runtime through an iterative proposal--evaluation loop to better solve target tasks. Existing meta-harness methods mainly adopt proposer-centric discovery, in which an LLM-based proposer integrates accumulated experimental findings to determine subsequent Harness revisions. This places the burden of maintaining the evolving improvement state on the proposer as history expands and its underlying experimental logic becomes harder to discern. In this paper, we introduce the Causal Improvement Graph (CIG), a graph-governed meta-harness framework that externalizes the evolving improvement state in a persistent graph, allowing prior findings to directly govern subsequent Harness optimization through local proposer operations. CIG grows and links Evidence, Hypothesis, Intervention, and Outcome nodes to represent what was observed, how it may be explained, how to test that explanation, and what the evaluation reveals. Their structural relations preserve how the improvement state changes across iterations, allowing local proposers to build directly on relations among prior findings rather than recover them from raw history. Across various agent tasks, CIG discovers stronger Harnesses than previous meta-harness baselines and remains robust to the choice of task solver and proposer. Structural ablations further support the design of an explicit improvement state with graph-governed evolution.