arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03324cs.CL

To Jev or Not? 评估用于仇恨言论审核的结构化决策模型的准确性与效率

To Jev or Not? Evaluating the Accuracy and Efficiency of Structured Decision Models for Hate-Speech Moderation

Demetris Paschalides, George Pallis, Marios D. Dikaiakos

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估结构化决策模型在仇恨言论审核中的准确性与效率,发现商业LLM仅在特定数据集上占优,而提供定义或分解问题对分类改善有限,但低成本决策模型接近商业性能。

中文摘要 AI 辅助

在线内容的规模使得仇恨言论审核具有挑战性,而大型语言模型(LLMs)使得有害材料的生成和改编更加容易。因此,审核需要能够适应不同仇恨言论定义的高效分类器。最近的结构化决策模型接受自然语言标准并在指定答案中进行选择,这引发了它们是否能在无需特定任务训练的情况下满足这些要求的问题。我们提出了HATEDECIDE,这是一个对四种仇恨言论数据集上的六种决策模型配置进行的评估,并与专门的审核模型、零样本模型、商业模型和监督基线进行比较。我们研究了提供数据集的定义,或将其分解为多个问题,是否能改善分类,并测量了它们的延迟和成本。我们发现,商业LLMs仅在一种数据集上显著优于所有决策模型。提供定义改变了高达28%的预测,但并未一致地改善分类,而分解仅在20%的比较中显著提升了性能。在一个诊断性测试用例集上,最佳托管的决策模型在约97%更低的推理成本下,与最佳商业LLM的宏F1分数相差1.6个百分点。这些结果指出了低成本审核的机会,同时表明明确的标准和额外的问题并不能可靠地改善分类。

英文摘要

The scale of online content makes hate-speech moderation challenging, while Large Language Models (LLMs) enable harmful material to be produced and adapted more easily. Moderation therefore requires efficient classifiers that can accommodate different definitions of hate speech. Recent structured decision models accept natural-language criteria and select among specified answers, raising the question of whether they can meet these requirements without task-specific training. We present HATEDECIDE, an evaluation of six decision-model configurations on four hate-speech datasets against specialized moderation, zero-shot, commercial, and supervised baselines. We examine whether supplying a dataset's definition, or decomposing it into multiple questions, improves classification, and we measure their latency and cost. We find that commercial LLMs significantly outperform all decision models on only one dataset. Supplying definitions changes up to 28\% of predictions without consistently improving classification, and decomposition significantly improves performance in only 20\% of the comparisons. On a diagnostic set of test cases, the best hosted decision model comes within 1.6 macro-F1 points of the best commercial LLM at approximately 97\% lower inference cost. These results identify opportunities for inexpensive moderation, while showing that explicit criteria and additional questions do not reliably improve classification.

发表机构

  • University of Cyprus(塞浦路斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑