arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体评估何时结束?结果终局性与跨单元分离

When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation

Avyay M. Casheekar, Hariganesh Tangirala

arXiv 2608.14940首次发表:更新:

发表机构

University of Michigan Law School(密歇根大学法学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有智能体评估将终点分数视为最终结果的问题,该研究提出结果终局性与跨单元分离两个条件,通过实验和协议审查指出当前评估的不足,并提出开放效应记录方案以完善评估。

AI 中文摘要

当前的智能体评估在已停止的运行结束时,基于可见状态对模型打分,并将其计为一次试验。然而,将该分数解读为最终结果需要两个条件,而终点本身未必能确立这两个条件:结果终局性与跨单元分离。这两个条件相互独立,因为协调延迟结果可确定标签,同时运行仍共享状态;而隔离运行可防止状态延续,同时待评分结果仍未完成。我们提出一种完成论证,明确每项决策所需的证据,并认为仅当所有可能改变所主张结果的事项均已解决、受限或作为不确定性保留时,最终标签才合理。首先,在受控重放中,为展示智能体动作固定时的机制,我们发现对于每项延迟操作,终点与终端标签均不同;当服务状态在运行间持续存在时,延迟写入会改变下一次运行的分数,但在隔离或经验证的重置后则不会。其次,在对10个公开协议的审查中,我们发现所有协议均明确运行何时停止及评分内容,而未完成操作与将运行视为独立试验的证据则记录得较不一致。最后,我们提出一种开放效应记录,列出终点后仍可能相关的操作或资源、其当前状态,以及它们是否可能改变待评分结果或影响其他运行。

英文摘要

Agent evaluations commonly score the state observed when a run stops and count the run as one trial. Interpreting that score as a final result from a separate trial requires outcome finality and cross-unit separation. Outcome finality requires that later events cannot change the claimed result, while cross-unit separation requires that earlier runs cannot change the relevant conditions of later ones. The endpoint establishes neither condition by itself, and the two can hold independently. Waiting for a delayed outcome may settle the label even though its state remains available to another run. Isolation may prevent carryover even though the scored outcome remains unresolved. We develop a completion argument that identifies the evidence needed for each decision. A final success or failure label is justified only when every relevant effect is resolved or bounded tightly enough to fix the outcome. Any remaining uncertainty must be reported. First, in a controlled replay with fixed agent actions, we find that endpoint and terminal labels differ for every nonzero-delay operation and that a delayed write changes the next run's score under shared state but has no such effect after namespacing or verified reset. Second, in a review of ten public protocols, we find that reset or deliberate retention is documented explicitly more often than unfinished operations or evidence for separate scoring. Finally, we propose an open-effects record for operations and resources that may remain relevant after the endpoint, their status, and their possible effects on the scored outcome or another run.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑