通过查询条件归因审计智能体行为
Auditing Agent Actions through Query-Conditioned Attribution
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- IBM(国际商业机器公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出查询条件智能体行为归因任务及A3Bench基准,利用小型开放权重模型结合梯度显著性与语义相关性高效归因,显著提升准确率并降低延迟。
AI中文摘要:
大语言模型智能体通过与用户、策略和外部工具的交互日益采取具有重大后果的行动。审计这些智能体需要将已实现的行为自动归因到其历史依据。然而,现有的归因公式无法为多样化的审计目标提供特定于查询的轨迹。此外,当对行为模型的访问受限时(例如,仅API部署),适用的方法通常依赖昂贵的输入扰动或外部LLM对完整轨迹的分析。因此,我们提出了查询条件智能体行为归因这一新任务,该任务以自然语言审计查询作为输入,并恢复查询指定行为方面的来源和有序中间证据。我们使用A3Bench实例化该任务,该基准包含1,396个审计查询,涵盖策略依据、参数来源、故障传播和不安全行为追踪。为了实现高效、特定于查询的归因,我们使用小型开放权重模型作为归因提议器,将查询条件梯度显著性查询语义相关性相结合,对历史单元进行排序。我们的提议器在较低推理成本下持续实现比开放权重基线更强的来源和证据排名,仅通过两次前向传播和一次反向传播,将来源MRR提升高达40.9%,证据MAP提升42.1%。受控评估证实,我们的提议器通过适应审计查询的细粒度变化来提高归因特异性。基于提议器集成,我们的端到端系统在来源准确性上超越了最强前沿模型基线(64.5%对60.4%),同时相对于最快前沿API基线将经验部署延迟降低了29.9%。代码和数据将在最终验证和清理后的初始审查期后发布。
英文摘要:
LLM agents increasingly take consequential actions through interactions with users, policies, and external tools. Auditing these agents requires automated attribution of realized actions to their historical basis. However, existing attribution formulations do not provide question-specific traces for diverse auditing objectives. Additionally, when access to the acting model is limited (e.g., in API-only deployments), applicable methods commonly rely on costly input perturbations or external LLM analysis of complete trajectories. We therefore formulate query-conditioned agent action attribution, a new task that takes a natural-language auditing query as input and recovers the source and ordered intermediate evidence for the query-specified aspect of an action. We instantiate this task with $A^3Bench$, a benchmark comprising 1,396 auditing queries across policy basis, parameter provenance, failure propagation, and unsafe-behavior tracing. To enable efficient, query-specific attribution, we use small open-weight models as attribution proposers that combine query-conditioned gradient saliency with query-semantic relevance to rank history units. Our proposer consistently achieves stronger source and evidence rankings at lower inference cost than open-weight baselines, improving source MRR by up to 40.9\% and evidence MAP by 42.1\% with only two forward passes and one backward pass. Controlled evaluations confirm that our proposer improves attribution specificity by adapting its rankings to fine-grained changes in the auditing query. Building on a proposer ensemble, our end-to-end system surpasses the strongest frontier-model baseline in source accuracy (64.5\% vs.\ 60.4\%) while reducing empirical deployment latency by 29.9\% relative to the fastest frontier API baseline. Code and data will be released after the initial review period following final validation and cleanup.