arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02925cs.AI

正未标注学习用于智能体安全误报审计

Positive-Unlabeled Learning for Agent Safety False Alarm Auditing

  • Jinan University(暨南大学)
  • Northwestern University(西北大学)
  • Analogy AI, Inc.(Analogy AI公司)
  • Shenzhen Technology University(深圳技术大学)
  • Tsinghua University(清华大学)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

Xichen Yan, Chongyang Gao, Kezhen Chen, Guangyi Zhang, Jiaqi Wu, Lixu Wang

中文总结 AI 辅助

针对语言模型智能体安全监控误报审计,提出正未标注学习框架,通过信任感知PU监督和可靠性门控排名蒸馏,无需警报标签,显著提升误报恢复率。

中文摘要 AI 辅助

安全监控器有助于保护与外部工具和环境交互的语言模型智能体,但保守的监控可能产生大量误报,消耗大量审查资源并削弱对警报的信任。由于误报和真实警报通常交织在原生监控分数中,获得可靠的截断值仍需要大量人工验证。在实践中,可能有一小部分经过验证安全的非警报轨迹可用,而警报仍处于未标注状态,这自然将误报审计视为一个正未标注(PU)排序问题。关键挑战是监控引起的选择偏差,因为观察到的安全参考被监控器接受,而感兴趣的被隐藏的安全误报恰恰是那些被监控器错误标记的,使得观察到的正样本对要恢复的正样本代表性不足。为解决这一挑战,我们提出了一个两阶段框架,其中信任感知的PU监督将安全参考适应到警报域,并保护可能的误报免受过度负压力,而可靠性门控的排名蒸馏将多个PU参考模型的一致排序偏好整合到单个学生模型中。共识引导的结构细化随后利用分层安全参考支持、警报关系和预测的参考共识来改进学生排名。该框架在拟合时不需要警报安全标签,并保持底层监控器不变。在主流安全监控器上,我们的方法实现了宏观AUPRC为0.6444,比八个评估的PU基线高出5.27至16.98个绝对百分点;与评估的最强PU基线PULDA相比,在5%的审查预算下多恢复了33.3%的误报。

英文摘要

Safety monitors help safeguard language-model agents interacting with external tools and environments, but conservative monitoring can generate many false alarms, consuming extensive review resources and weakening trust in alerts. Because false and genuine alarms often remain interleaved in native monitor scores, obtaining a reliable cutoff still requires substantial manual verification. In practice, a small set of verified-safe non-alarmed trajectories may be available while alarms remain unlabeled, naturally casting false-alarm auditing as a positive-unlabeled (PU) ranking problem. The key challenge is monitor-induced selection, since observed safe references are accepted by the monitor, while the hidden safe alarms of interest are precisely those it incorrectly flags, making the observed positives poorly representative of the positives to be recovered. To address this challenge, we propose a two-stage framework in which Trust-aware PU Supervision adapts safe references toward the alarm domain and protects plausible false alarms from excessive negative pressure, while Reliability-gated Rank Distillation consolidates consistent ordering preferences from multiple PU reference models into a single student. Consensus-guided Structural Refinement then improves the student ranking using hierarchical safe-reference support, alarm relations, and predicted reference consensus. The framework requires no alarm safety labels for fitting and leaves the underlying monitor unchanged. Across mainstream safety monitors, our method achieves a macro AUPRC of $0.6444$, outperforming eight evaluated PU baselines by 5.27--16.98 absolute percentage points; compared with PULDA, the strongest evaluated PU baseline, it recovers 33.3% more false alarms at a 5% review budget.

补充信息

↑