arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03814cs.CL

评估大型语言模型在内容审核中的条件准则性行为

Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation

  • University of Sheffield(谢菲尔德大学)

机构由 AI 辅助整理,请以论文原文为准。

Danting Zhang, Bei Peng, Robert Loftin

AI总结:

本研究针对内容审核场景,提出DECO工具与成对评估方法,发现大型语言模型在聚合标签基准上表现优异却存在准则级缺陷,需开发针对性评估方法。

AI中文摘要:

大型语言模型(LLMs)在标准内容审核基准上表现出较强性能。然而这些基准通常将多个审核准则聚合为单一标签,导致模型是否能区分各准则并在决策时可靠应用每个准则尚不明确。为研究LLMs是否表现出条件准则性行为,我们引入了内容诊断评估工具(DECO),这是一种与准则无关的内容分解方法,可实现受控的准则级评估。我们还引入了成对评估方法,用于比较同一输入在不同准则下的模型输出。在四个审核数据集和四个LLMs上,我们发现优异的基准性能可能掩盖准则层面的重大失败。当正确决策不取决于整体危害性,而是取决于准则要求评估的内容特定方面时,模型表现最差。我们的结果凸显了当前内容审核基准的一个关键局限:聚合标签上的优异性能不足以证明LLMs能可靠地针对单个审核准则评估内容。这些发现呼吁开发能明确测量条件准则性行为的评估方法。

英文摘要:

Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However, these benchmarks often aggregate multiple moderation criteria into a single label, making it unclear whether models can disentangle them and reliably apply each criterion when making decisions. To study whether LLMs exhibit criterion-conditioned behaviour, we introduce Diagnostic Evaluation of COntent (DECO), a criterion-independent factorisation of content that enables controlled, criterion-level evaluation. We also introduce pairwise evaluation to compare model outputs across different criteria for the same input. Across four moderation datasets and four LLMs, we find that strong benchmark performance can hide substantial failures at the criterion level. Models struggle most when correct decisions depend not on overall harmfulness, but on the specific aspect of the content that the criterion requires them to assess. Our results highlight a key limitation of current content moderation benchmarks: strong performance on aggregated labels does not provide sufficient evidence that LLMs can reliably evaluate content with respect to individual moderation criteria. These findings call for the development of evaluation methods that explicitly measure criterion-conditioned behaviour.

补充信息

↑