arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谁来评判评判者?用于评估大语言模型响应与安全评判者的中文安全问答基准

Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges

Rui Yang, Shuang Huang, Junhua Liu, Ziqi Zhao, Qingzhong Yan, Yuhang Sun, Cong Liu, Guoping Hu, Rui Mei, Jing Shao

arXiv 2609.01210首次发表:更新:

发表机构

Anhui SparkShield Intelligent Technology; iFLYTEK; Peking University; Shanghai Innovation Institute; Shanghai Artificial Intelligence Laboratory(安徽星火盾智能科技; 讯飞科技; 北京大学; 上海创新研究院; 上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出中文安全问答基准C-SafeQA,基于多模型裁决与专家盲审生成参考标签,用于评估大语言模型响应安全及自动化安全评判者,发现评判者在关键指标间存在权衡,藏头诗转换会降低评判者的不安全响应召回率。

AI 中文摘要

大语言模型的安全基准通常评估用户查询的风险,尽管问答结果取决于响应是否违反政策。这种区分在中文有害内容评估中至关重要,因为语言变异和对抗性转换可能掩盖危险意图。我们推出C-SafeQA,这是一个基于政策的响应级中文安全评估基准,包含538个基础查询和8877个对抗性查询,由四个全模型大语言模型部署给出答案,产生37660条标记为安全、不安全或有争议的查询-响应记录。参考标签通过感知一致性的多模型裁决,以及三名安全专家对分层子集的盲审生成。C-SafeQA既支持目标模型的安全评估,也支持针对共享参考标签对七个自动化安全评判者进行审计。基础查询的不安全响应率在0.93%至3.35%之间,对抗性查询的不安全响应率在11.68%至30.05%之间。在对抗性子集上,评判者在不安全响应召回率和风险查询条件下的安全响应误报率之间存在显著权衡,且没有任何评判者在所有指标上占优。藏头诗转换会降低所有七个评判者的不安全召回率,暴露出特定机制的评估者弱点。数据集记录、元数据、验证代码和评判者脚本已公开发布以支持复现,而基准构建、目标响应生成和私人裁决仍不在发布范围内。

英文摘要

Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the response violates a policy. This distinction is critical in Chinese harmful-content evaluation, where linguistic variation and adversarial transformations can obscure risky intent. We introduce C-SafeQA, a policy-grounded benchmark for response-level Chinese safety evaluation. It comprises 538 base queries and 8,877 adversarial queries answered by four full-model LLM deployments, yielding 37,660 query-response records labeled safe, unsafe, or disputed. Reference labels are generated through agreement-aware multi-model adjudication and blind audits of stratified subsets by three safety experts. C-SafeQA supports both evaluation of target-model safety and auditing of seven automated safety judges against shared reference labels. Unsafe-response rates range from 0.93% to 3.35% on base queries and from 11.68% to 30.05% on adversarial queries. On the adversarial subset, judges show substantial trade-offs between unsafe-response recall and risk-query-conditioned safe-response false positive rate, and no judge dominates all metrics. Both acrostic transformations reduce unsafe recall for all seven judges, revealing mechanism-specific evaluator weaknesses. Dataset records, metadata, verification code, and judge scripts are publicly released to support recomputation, while benchmark construction, target-response generation, and private adjudication remain outside the release boundary.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑