arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07089cs.CRcs.AIcs.LG

迈向统一的滥用监控基准

Towards a Unified Misuse Monitoring Benchmark

Aniruddh Pramod, James Oldfield, Adel Bibi

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM智能体面临的分解攻击和提示注入攻击,提出统一的滥用监控框架,通过监控行动识别伤害窗口,构建约6200条对话基准,行动框架监控器在经典指标下表现优异(AUC 0.95/0.99),并揭示位置盲指标的乐观偏差。

中文摘要 AI 辅助

LLM智能体日益在多参与者环境中行动,使其面临来自多种来源的滥用:分解攻击(将有害请求拆分为无害的子请求)和提示注入攻击(被入侵的工具传递恶意指令)。现有评估将这些威胁分开处理,并询问轨迹是否有害,而非何时变得有害。我们提出监控智能体的响应(其行动在此被外部化),并询问监控器识别出伤害的第一个点是否落在伤害窗口内(从智能体首次做出有害承诺到目标执行)。我们开发了一种用于轨迹级滥用监控的统一形式化方法,并利用它构建了一个包含约6,200个用户、LLM智能体和外部环境之间的对话记录基准,覆盖了共享模式中的两种威胁,带有标记的伤害窗口、相应的良性对照以及对这些请求的拒绝匹配实例。在17种监控器配置中,我们发现我们提出的行动框架监控器在经典指标下对两种威胁均表现良好(AUC分别为0.95和0.99),而内容框架监控器在注入攻击上崩溃(AUC为0.52)。我们还表明,经典的位置盲指标对监控器性能描绘了乐观图景,因为所有监控器在区间指标(衡量定位伤害的能力)下对分解攻击的定位能力都很差。总体而言,我们说明了对滥用监控进行统一研究的必要性。

英文摘要

LLM agents increasingly act in multi-actor environments, exposing them to misuse from multiple sources: decomposition attacks, where a harmful request is split into innocuous sub-requests, and prompt injection attacks, where a compromised tool delivers a malicious instruction. Existing evaluations treat these threats separately and ask whether a trajectory is harmful, rather than when it becomes harmful. We propose monitoring the agent's responses, where its actions are externalised, and ask whether the first point where monitors identify harm lands within a harm window (from the agent's first harmful commitment to goal execution). We develop a unified formalism for trace-level misuse monitoring and use it to construct a benchmark of ~6,200 conversation transcripts between a user, an LLM agent, and the external environment, spanning both threats in a shared schema, with a labelled harm window, corresponding benign controls, and matched instances of refusals to these requests. Across 17 monitor configurations, we find that our proposed action-framed monitors perform well on both threats under classical metrics (AUC: 0.95 and 0.99 respectively), while content-framed monitors collapse on injection attacks (AUC: 0.52). We also show that classical position-blind metrics paint an optimistic picture of monitor performance, since all monitors localise decomposition attacks poorly under the interval metric, which measures the ability to localise harm. Broadly, we illustrate the need for a unified study of misuse monitoring.

发表机构

  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑