arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过合作-义务耦合对大语言模型智能体(LLM-agent)的涌现协作进行审计

Auditing Emergent LLM-Agent Collaboration through Cooperation-Obligation Coupling

Zuyuan Zhang, Hanqing Yang, Carlee Joe-Wong, Tian Lan

arXiv 2607.27429首次发表:更新:

AI 中文总结

针对LLM-agent协作审计的可审计性缺口,本文提出iCORE表示方法,建立合作-义务耦合审计框架,实验验证其可准确检测缺陷并显著提升轨迹与终端性能。

AI 中文摘要

LLM-agent系统可通过动态自组织和涌现协作解决复杂任务,对该过程进行审计至关重要,因为看似合理的中间或最终输出可能掩盖不完整或无支撑的工作、责任分配不当,最终损害响应质量。现有方法可能记录消息、工具调用、来源或任务依赖,但存在可审计性缺口——它们未联合表示剩余工作内容、负责主体及各工作状态转换的证据依据。针对该缺口,本文提出Integrated Cooperation-Obligation REpresentation(iCORE,即集成合作-义务表示),创建统一编码$X=(G,Q,\boldsymbol{\rm \textit{\textPi}})$,将可观测交互整合为合作图$G$,将动态工作与任务分配整合为义务图$Q$,并将二者与可验证属性及证据关联的审计映射$\boldsymbol{\rm \textit{\textPi}}$。该iCORE表示使审计者可验证两个互补属性:一是工作合理性,即每项与决策相关的有效工作断言必须通过$G$和$\boldsymbol{\rm \textit{\textPi}}$获得有限依据;二是智能体分配稳定性,即不存在可行替代智能体,能为评估的义务将声明贡献值提升超过$\boldsymbol{\rm \textit{\textepsilon}}$。本文在给定条件下建立了局部到全局的合理性、分配遗憾保证及性能边界。iCORE是工作流之上的检测层,数值结果显示,完整耦合状态可准确重构两种执行模式下的合理性与分配缺陷;相较于被动全状态观察,iCORE-Audit在受控执行和真实LLM执行中,分别实现11.5%和26.4%的绝对轨迹质量提升,对应绝对终端性能提升为15.1%和31.0%。

英文摘要

LLM-agent systems can solve complex tasks through dynamic self-organization and emergent cooperation. Auditing this process is essential because plausible intermediate or final outputs can conceal incomplete or unsupported work and poorly allocated responsibility, ultimately compromising response quality. While existing approaches may record messages, tool calls, provenance, or task dependencies, an auditability gap exists as they do not jointly represent what work remains, who is responsible for it, and what evidence justifies each work-state transition. We address this auditability gap by proposing \emph{Integrated Cooperation-Obligation REpresentation} (iCORE). It creates a unified encoding $X=(G,Q,Π)$ integrating observable interactions as a cooperation graph $G$, evolving work and assignments as an obligation graph $Q$, and the audit map $Π$ linking them with verifiable properties and evidence. This iCORE representation enables the auditor to certify two complementary properties: {Work soundness}, where every active decision-relevant work assertion must have a finite justification through $G$ and $Π$; and {Agent-assignment stability}, which requires that no feasible alternative agent improve the declared contribution value for an evaluated obligation by more than $ε$. We establish local-to-global soundness and assignment-regret guarantees and a performance bound under stated conditions. iCORE is an instrumentation layer over workflows. Numerical results show that the full coupled state exactly reconstructs soundness and assignment defects in two execution modes and that, relative to passive full-state observation, iCORE-Audit yields absolute trajectory-quality improvements of $11.5\%$ and $26.4\%$ in controlled and real-LLM execution, respectively, with corresponding absolute terminal-performance improvements of $15.1\%$ and $31.0\%$.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑