arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04575cs.CLcs.AI

从探针分数到告警策略:语言模型智能体激活监控器的操作有效性

From Probe Scores to Alarm Policies: Operational Validity of Activation Monitors for Language-Model Agents

Xueping Gao

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出操作有效性契约,评估激活监控器在告警策略下的实际表现,发现其AUROC高但锁定阈值下检测率低,无法达到策略有效性。

中文摘要 AI 辅助

激活探针能以高接收者操作特征曲线下面积(AUROC)预测语言模型的安全相关属性,但部署的智能体监控器在严格的误报预算下做出阈值化告警决策。这些是不同的估计目标。我们引入操作有效性契约,该契约固定监控器的目标、可观测性、身份、时序、干预单元、比较器、校准和成本。我们在语义请求或轨迹层面形式化风险:当一个任务包含重复的告警机会时,行级AUROC和假阳性率无法识别语义单元的任何告警风险。置信认证阈值还需要足够的独立负单元、平局安全规则以及部署时的迁移。在Models Under Pressure、LASR拒绝预测、不可变AgentDojo以及前瞻性协议冻结的ST-WebAgentBench复制中,联合激活可观测监控器在前两个基准上达到AUROC 0.957和0.935,但其锁定的5%/10%检测率分别仅为.642/.742和.719/.782。在AgentDojo上,次级平均激活展开监控器达到AUROC 0.922,但在锁定的5%操作点未检测到38个阳性语义案例中的任何一个;旨在10%误报的阈值在测试中实现18.7-20.0%。由于测试未达到其前瞻性冻结的40个阳性支持门控,我们将其标记为支持不足。在ST-WebAgentBench上,激活达到AUROC .874,但23个独立校准阴性无法识别即使10%的控制者;锁定策略弃权(不执行),而非将其机械零FPR报告为成功。一个探索性反例还降低了完整AUROC,同时提高了实现的10%效用。故障关闭编译器将MUP和LASR限制在受限预测值,AgentDojo和ST-Web限制在表示可访问性;没有设置达到告警策略有效性。

英文摘要

Activation probes can predict safety-relevant properties of language models with high area under the receiver-operating-characteristic curve (AUROC), but deployed agent monitors make thresholded alarm decisions under tight false-alarm budgets. These are different estimands. We introduce an Operational Validity Contract that fixes a monitor's target, observability, identity, timing, intervention unit, comparator, calibration, and cost. We formalize risk at the semantic request or trajectory level: when one task contains repeated alarm opportunities, row-level AUROC and false positive rate do not identify semantic-unit any-alarm risk. A confidence-certified threshold also requires enough independent negative units, a tie-safe rule, and transport to deployment. Across Models Under Pressure, LASR refusal prediction, immutable AgentDojo, and a prospectively protocol-frozen ST-WebAgentBench replication, joint activation-observable monitors reach AUROC 0.957 and 0.935 on the first two benchmarks, yet their locked 5%/10% detection rates are only .642/.742 and .719/.782, respectively. On AgentDojo, the secondary mean-activation rollout monitor reaches AUROC 0.922 but detects none of 38 positive semantic cases at the locked 5% operating point; thresholds intended for 10% false alarms realize 18.7-20.0% on test. Because the test misses its prospectively frozen 40-positive support gate, we label it support-insufficient. On ST-WebAgentBench, activation reaches AUROC .874, but 23 independent calibration negatives cannot identify even a 10% controller; the locked policy abstains rather than reporting its mechanical zero FPR as a success. An exploratory counterexample also lowers full AUROC while improving realized 10% utility. The fail-closed compiler caps MUP and LASR at restricted predictive value and AgentDojo and ST-Web at representation accessibility; no setting reaches alarm-policy validity.

发表机构

  • Alibaba Cloud Computing(阿里云计算)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑