arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VARM-Bench:中文辱骂内容审核中可验证的结构化推理基准测试

VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation

Mingyu Yuan, Shengtao Wen, Lingbing Guo, Zhen Bi, Xiang Chen

arXiv 2608.15600首次发表:更新:

发表机构

NUAA(南京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对中文辱骂内容审核缺乏可验证决策依据的问题,提出VARM-Bench基准,通过确定性协议评估模型,发现标签级性能可能掩盖记录错误,提供可审计的评估基准。

AI 中文摘要

网络辱骂内容的广泛传播,使得对中文社交媒体文本进行可靠审核的需求日益增长。现有中文基准测试支持标签分类、细粒度毒性分类和目标感知提取,但未提供统一表示以确定性验证审核决策的依据。我们推出VARM-Bench,这是一个针对中文辱骂内容审核的领域锚定思维链理由基准测试。每个实例包含带有明确锚点的简洁自然语言理由,对应六个决策:目标、目标类型、目标明确性、作者立场、危害性标签和细粒度类别。我们的确定性协议评估领域正确性、目标对齐、输出有效性、完整记录一致性,以及在最终决策正确条件下的隐藏记录错误,且不依赖大语言模型(LLM)评判。在通用结构化输出协议下,我们使用零样本提示、分类法指导和结构化思维链监督,评估多个模型家族的语言模型,并分析词汇线索敏感性和领域级错误。结果显示,强大的标签级性能可能掩盖完整审核记录中的大量错误。VARM-Bench为评估中文辱骂内容审核中可验证的审核理由提供了可审计且可复现的基准测试。

英文摘要

The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text. Existing Chinese benchmarks support label classification, fine-grained toxicity categorization, and target-aware extraction, but do not provide a unified representation for deterministically verifying the stated basis of a moderation decision. We introduce VARM-Bench, a benchmark for field-anchored chain-of-thought rationales in Chinese abusive-speech moderation. Each instance contains a concise natural-language rationale with explicit anchors for six decisions: target, target type, target explicitness, author stance, harmfulness label, and fine-grained category. Our deterministic protocol evaluates field correctness, target alignment, output validity, complete-record agreement, and hidden record errors conditioned on correct final decisions, without relying on an LLM judge. Under a common structured-output protocol, we evaluate language models across multiple model families using zero-shot prompting, taxonomy guidance, and structured CoT supervision, and analyze lexical-cue sensitivity and field-level errors. Results show that strong label-level performance can conceal substantial errors in complete moderation records. VARM-Bench provides an auditable and reproducible benchmark for evaluating verifiable moderation rationales in Chinese abusive-speech moderation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑