arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

审计LLM智能体环境中的动作结算:顺序、进度与重放

Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay

Haotian Chen, Bowen Ye, Yuning Zhang, Jingkun Yu

arXiv 2610.01138首次发表:更新:

发表机构

University of Science and Technology of China; Shanghai Jiao Tong University; Southwest Jiaotong University(中国科学技术大学; 上海交通大学; 西南交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过类型化快照-结算契约审计LLM智能体环境的动作结算,测试五种策略,发现保守拒绝完成率低,而优先级仲裁非最优,证据聚焦执行语义。

AI 中文摘要

在大语言模型(LLM)智能体环境中,并发动作即使每个提议单独有效,也需要仲裁。我们实现了一个类型化的快照-结算契约,并审计了三个不同的属性:顺序敏感性、有效进度和重放一致性。在28,800次穷举排列试验和2,160个脚本化多步情节中测试了五种结算策略。联合策略在固定优先级条件下具有空间顺序不变性,但保守拒绝在六智能体门口任务中仅完成31.25%的智能体,而随机票证完成率为90.28%;配对改进为59.03个百分点(95%自助法区间:50.00-68.06)。所有策略均保持所测试的空间约束,且优先级仲裁仍错过独立小实例最优解。一个单独的全状态日志审计精确重放了156个检查点,并拒绝了1,332个构造的损坏,且保留了终端锚点。证据涉及执行语义,而非人类真实性或长期公平性。

英文摘要

Concurrent actions in large language model (LLM) agent environments require arbitration even when each proposal is individually valid. We implement a typed snapshot-settlement contract and audit three distinct properties: order sensitivity, useful progress, and replay consistency. Five settlement policies are tested in 28,800 exhaustive permutation trials and 2,160 scripted multistep episodes. Joint policies are spatially order-invariant conditional on fixed priorities, yet conservative rejection completes only 31.25% of agents in a six-agent doorway task versus 90.28% for random tickets; the paired improvement is 59.03 percentage points (95% bootstrap interval: 50.00-68.06). All policies preserve the tested spatial constraints, and priority arbitration still misses the independent small-instance optimum. A separate full-state journal audit exactly replays 156 checkpoints and rejects 1,332 constructed corruptions with a retained terminal anchor. The evidence concerns execution semantics, not human realism or long-run fairness.

Comments5 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑