AI 中文总结
该研究以400本受限与非受限书籍为测试平台,对6家AI提供商的6种前沿模型开展大规模实验,发现现代LLM对受限书籍仅0.07%的案例拒绝对话,其内容政策已转向校准的上下文敏感披露。
AI 中文摘要
随着大语言模型(LLM)进入日常信息流通渠道,理解它们如何处理敏感话题与理解它们是否处理这些话题同样重要。我们通过大规模系统性实验,以受限书籍与非受限书籍作为受控测试平台来研究该问题:共40800组查询-响应对、涉及400本书籍、17种提示设计,以及来自6家AI提供商的6种前沿模型(Claude Sonnet 4.5、GPT-4o、Gemini 2.5 Flash、DeepSeek-V3、Qwen-Plus和Grok-4.1-Fast)。我们的受限书籍集源自美国图书馆协会(ALA)2000-2023年最受质疑书籍记录;全程使用“受限”而非“ banned”,因为ALA文件记录的是正式挑战——要求移除或限制访问,并非都导致彻底封禁。我们的核心发现是“零弃权(不执行)”现象:现代LLM仅在0.07%的案例中拒绝对受限书籍进行讨论,这实际上使针对此类内容类别的越狱研究的前提失效。差异体现在警告语言(占比提升8-15个百分点,p<0.001)和犹豫标记(占比提升2-5个百分点)上,其中性内容提及率是最强的个体信号(占比提升33-52个百分点)。我们还识别出不同提供商之间的系统性差异,并表明仅提示框架就可使警告率差距变化达19个百分点。这些结果表明,LLM的内容政策已从二元拒绝对话转向校准的、上下文敏感的披露——该发现在西方和中国AI提供商中均一致成立。
英文摘要
As large language models enter everyday information pipelines, understanding how they handle sensitive topics matters as much as understanding whether they handle them at all. We study this question through a large-scale, systematic experiment using restricted versus unrestricted books as a controlled testbed: 40,800 query-response pairs, 400 books, 17 prompt designs, and six frontier models spanning six AI providers (Claude Sonnet 4.5, GPT-4o, Gemini 2.5 Flash, DeepSeek-V3, Qwen-Plus, and Grok-4.1-Fast). Our restricted set is drawn from the American Library Association's Most Challenged Books records (2000-2023); we use restricted rather than banned throughout because the ALA documents formal challenges-requests to remove or restrict access-which do not always result in outright bans. Our central finding is a zero-refusal phenomenon: modern LLMs decline to discuss restricted books in only 0.07% of cases, effectively invalidating the premise of jailbreaking research for this content class. Differentiation occurs instead through warning language (+8-15 percentage points, p < 0.001) and hesitation markers (+2-5 pp), with sexual content mention rate as the strongest individual signal (+33-52 pp). We further identify systematic differences between providers and show that prompt framing alone shifts the warning-rate gap by up to 19 pp. These results indicate that LLM content policy has shifted from binary refusal toward calibrated, context-sensitive disclosure-a finding that holds consistently across Western and Chinese AI providers.
CommentsAccepted at AIES 2026 (AAAI/ACM Conference on AI, Ethics, and Society). 13 pages, 4 figures, 5 tables. This arXiv version includes the full appendix