arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Safety-Flag:LLM 内容审核器可靠性与校准的统一基准

Safety-Flag: A Unified Benchmark for the Reliability and Calibration of LLM Content Moderators

Yibo Hu

arXiv 2609.19072首次发表:更新:

发表机构

Illinois Institute of Technology(伊利诺伊理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Safety-Flag 是一个统一基准,整合七个安全基准,评估 LLM 内容审核器的错误方向、校准和置信度排序,发现总体准确率掩盖严重偏差,温度校准可大幅降低误差。

AI 中文摘要

大型语言模型越来越多地用于内容审核,但大多数评估仍然报告单个基准上的总体准确率。我们引入了 Safety-Flag,它将七个广泛使用的安全基准(BeaverTails、XSTest、Ethics、WildGuard、Aegis、ToxiChat 和 ToxiGen)整合到一个平衡的标记/不标记协议中。我们发布了六个通用大语言模型和四个专用审核模型,以及三个参考模型,在相同项目上的逐项决策和置信度分数。Safety-Flag 衡量审核器可靠性的三个维度:错误方向、概率校准以及用于人工审核的基于置信度的错误排序。这些维度常常相互矛盾。总体准确率无法揭示错误方向:一个模型标记了 85% 的无害内容,而另一个模型漏掉了 54% 的有害内容。所有六个通用模型都过度自信;为每个模型拟合一个温度参数可将校准误差降低 2.8 至 6.0 倍,而不改变预测标签或置信度排序。基于置信度的弃权(不执行)降低了每个模型的选择性风险,尽管收益取决于置信度对错误排序的好坏。专用审核模型产生更少的误报且校准更好,但有几个在其文档覆盖范围之外具有更高的漏报率。我们在以下网址发布了基准、固定项目列表、评估代码、逐项模型输出和排行榜:此 https URL。

英文摘要

Large language models are increasingly used for content moderation, but most evaluations still report aggregate accuracy on individual benchmarks. We introduce Safety-Flag, which places seven widely used safety benchmarks (BeaverTails, XSTest, Ethics, WildGuard, Aegis, ToxiChat, and ToxiGen) into a single balanced flag / do-not-flag protocol. We release item-level decisions and confidence scores for six general-purpose LLMs and four dedicated guards, together with three reference models, evaluated on the same items. Safety-Flag measures three dimensions of moderator reliability: error direction, probability calibration, and confidence-based error ranking for human review. They often disagree. Aggregate accuracy does not reveal error direction: one model flags $85\%$ of benign content, whereas another misses $54\%$ of harmful content. All six general-purpose models are overconfident; fitting one temperature per model reduces calibration error by $2.8$--$6.0\times$ without changing predicted labels or confidence ordering. Confidence-based abstention lowers selective risk for every model, although the gains depend on how well confidence ranks errors. Dedicated guards produce fewer false alarms and are better calibrated, but several have higher miss rates outside their documented coverage. We release the benchmark, fixed item lists, evaluation code, per-item model outputs, and leaderboard at: https://github.com/yibo-hu-lab/safety-flag-benchmark.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑