arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32677cs.AIcs.LGcs.MA

当更好变得更糟:自适应世界中自我改进智能体的改进保真度

When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds

Ke Wang, Zijie Zhao, Zhiyi Yuan, Changlun Li

AI总结:

本研究提出改进保真度概念,揭示代理验证器在自适应世界中可能误导自我改进,并引入PIVOT-KG验证器,通过决策感知的评估分配显著降低选择遗憾,确保可靠改进。

AI中文摘要:

自我改进智能体日益依赖代理验证器来选择策略更新,然而部署可能会改变这些更新被评估的世界。因此,即使验证器在整体上对策略排序良好,一个在验证器看来更好的更新在部署后也可能变得更糟。我们将这一差距形式化为改进保真度,它询问代理改进是否在改进过程实际提出的更新上保持了部署改进的符号和排序。我们表明,全局策略准确性并不必然保证更新保真度:操作者转移和部署响应可能产生更新级错误,而候选边际则决定这些错误是否改变替换决策。我们引入了PIVOT-KG,一种成对的、决策感知的验证器,它根据每单位成本的预期选择遗憾减少来分配稀缺的高保真度评估。在Leduc、Kuhn和Melting Pot中的90个保留根上,代理和部署最优集在51个案例中是不相交的。在八候选HighwayEnv压力测试中,PIVOT-KG在主要预算下将平均改进选择遗憾从精确Uniform验证规则下的0.0435降低到0.0055。这些结果共同表明,可靠的自我改进应该在它们所诱导的世界中评估提议的改进,同时为在部署证据可能影响替换决策时分配稀缺部署证据提供了实用规则。

英文摘要:

Self-improving agents increasingly rely on proxy verifiers to choose policy updates, yet deployment can change the world in which those updates are evaluated. An update that looks better to the verifier can therefore become worse after deployment even when the verifier ranks policies well overall. We formalize this gap as Improvement Fidelity, which asks whether proxy improvements preserve the sign and ordering of deployment improvements over the updates an improvement process actually proposes. We show that global policy accuracy need not guarantee update fidelity: operator shift and deployment response can create update-level errors, while candidate margins determine whether those errors change the replacement decision. We introduce PIVOT-KG, a paired, decision-aware validator that allocates scarce high-fidelity evaluation according to the expected reduction in selection regret per unit cost. Across 90 held-out roots in Leduc, Kuhn, and Melting Pot, proxy and deployment optimal sets are disjoint in 51 cases. In an eight-candidate HighwayEnv stress test, PIVOT-KG reduces mean improvement-selection regret from 0.0435 under the exact Uniform validation rule to 0.0055 at the primary budget. Together, these results show why reliable self-improvement should evaluate proposed improvements in the worlds they induce, while providing a practical rule for allocating scarce deployment evidence when it can affect the replacement decision.

补充信息

↑