发表机构
Tongji University(同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对中国AI生成内容法规,设计了一个包含2303个问题、六个维度的框架,评估了20个LLM的合规性,发现国际模型合规性高,差异主要源于意识形态对齐维度,并建立了全球统一的监管基准。
AI 中文摘要
LLM的广泛采用导致了内容合规风险的不断升级。先前的工作致力于在英语语境中解决这些风险,但低估了中文内容的复杂性。本文遵循中国当前的AI生成内容合规要求,对20个知名LLM提供了评估结果,深入洞察中国的监管格局。我们设计了一个新颖的框架,通过2303个问题(涵盖六个不同维度,包括203个自构建的宪法问题)来评估合规率和拒绝率。该框架采用多个评判者,基于其层级对齐记忆独立生成裁决。我们的研究结果表明,即使使用标准中文问题,国际模型也表现出高水平的合规性,主要差异可能源于与意识形态对齐密切相关的维度。我们建立了一个监管基准,使全球AI社区能够在统一的法律依据合规要求下评估中文和非中文LLM。
英文摘要
The widespread adoption of LLMs has led to escalating content compliance risks. Prior works have contributed to addressing these risks in the English context, downplaying the complexity of Chinese language content. This paper follows China's current AI-Generated content compliance requirements and provides evaluation results on 20 notable LLMs, offering insight into China's regulatory landscape. We design a novel framework to assess the compliance and refusal rates with 2303 questions spanning six distinct dimensions, including 203 self-constructed constitutional questions. The framework employs several judges to generate verdicts independently based on their hierarchical alignment memory. Our findings show that international models also exhibit high levels of compliance despite the use of standard Chinese questions, and the main differences may stem from dimensions closely related to ideological alignment. We establish a regulatory benchmark that enables the global AI community to evaluate both Chinese and non-Chinese LLMs under a unified set of legally grounded compliance requirements.
Comments5 pages, 3 figures, with appendix still improving