arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

提升中国青少年大语言模型安全性的标准:一个基于文化的细粒度基准

Raising the Bar for Chinese Adolescent LLM Safety: A Culturally-Grounded, Fine-Grained Benchmark

Jinxiang Wang, Yifan Liu, Jing Tan, Xiangyu Zhao, Xin Yao, Xuetao Wei

arXiv 2609.35902首次发表:更新:

发表机构

Southern University of Science and Technology; City University of Hong Kong; Lingnan University(南方科技大学; 香港城市大学; 岭南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对中文青少年对话安全,提出基于中国文化的细粒度基准QH-Bench,含单轮与多轮测试,评估发现模型在青少年线下接触风险及关系压力下安全边界维持方面存在明显弱点。

AI 中文摘要

与青少年对话中的安全风险并不总是显而易见的。除非模型考虑用户的年龄、处境和之前的对话轮次,否则一个请求可能看似无害。现有的中文安全基准主要针对一般用户,对青少年安全的关注有限。单轮测试也遗漏了多轮对话中出现的风险。QH-Bench是一个面向青少年内容安全的中文基准,其场景基于中国社会和文化背景。单轮轨道包含715个测试项,组织为10个风险域、50个子域和143个细粒度风险场景。多轮轨道包含100个四轮轨迹,采用平衡的10×10设计,将相同的十个域与十种跨轮机制相结合。两个轨道使用相同的五级安全-帮助性量表和自动评判器,并具有特定于轨道的标准。对13个开放权重模型的评估发现,线下接触场景是一个共同的弱点。每个模型在与在线联系人、陌生群体和成人进行线下会面相关的超过一半的项目上获得负面分数。这包括单轮领先者InternLM2.5-20B;负面分数表示部分或明确促进风险的回答。多轮领先者GLM-4-32B在用户先建立关系再援引忠诚或保密性的完整轨迹中,有60%获得负面分数。这些发现为被评估的模型确定了两个优先事项:处理青少年线下接触风险和在与关系压力下维持安全边界。领先的总体分数并不表明这些具体弱点已被解决。

英文摘要

Safety risks in conversations with adolescents are not always explicit. A request may appear harmless unless a model considers the user's age, circumstances, and earlier turns. Existing Chinese safety benchmarks mainly target general users and give limited attention to adolescent safety. Single-turn tests also miss risks that emerge over several turns. QH-Bench is a Chinese-language benchmark for adolescent content safety, with scenarios grounded in Chinese social and cultural settings. The single-turn track contains 715 test items organized into 10 risk domains, 50 subdomains, and 143 fine-grained risk scenarios. The multi-turn track contains 100 four-turn trajectories in a balanced 10-by-10 design that combines the same ten domains with ten cross-turn mechanisms. Both tracks use the same five-level safety-helpfulness scale and automatic judge, with track-specific criteria. Evaluation of 13 open-weight models identifies offline-contact scenarios as a shared weakness. Every model receives negative scores on more than half of the items involving offline meetings with online contacts, unfamiliar groups, and adults. This includes InternLM2.5-20B, the single-turn leader; negative scores indicate responses that partially or clearly facilitate risk. GLM-4-32B, the multi-turn leader, receives negative scores on 60% of complete trajectories in which users build relationships before invoking loyalty or confidentiality. These findings identify two priorities for the evaluated models: handling adolescent offline-contact risks and maintaining safety boundaries under relational pressure. Leading aggregate scores do not establish that these specific weaknesses have been resolved.

Comments23 pages, 6 figures. Code and benchmark: https://github.com/WEILaboratory/QH-Bench

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑