arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07953cs.SIcs.CY

在线审核中的System One模型基准测试

Benchmarking System One Models in Online Moderation

Federico Mazzoni, Andrea Failla

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估System One模型在在线审核中的表现,发现Jev在多个基准上竞争力强且受益于检索先例,而Laya一致性较差,总体表明SOM驱动的审核前景广阔。

中文摘要 AI 辅助

在线审核系统必须应用不断变化的平台政策、社区规则和先前决策,同时产生可审计并可转交人工审核的决策。我们评估了System One模型(接受自然语言上下文但返回类型化选择、概率或分数的模型)是否能够支持这一场景。在五个审核基准测试中,Jev与专门的参考系统相比具有竞争力,在多个基于政策和有害内容的设置中达到或超过它们。然后,我们使用受控信息条件来分离书面规则、检索到的先例以及候选答案空间的限制。Jev通常受益于检索到的先例,在答案空间保持固定时,提高了精确政策选择和有害内容判别能力。Laya的一致性较差:检索常常改变其正向预测率或无违规率,而未改善判别能力。Jev的置信度可以在多种设置中支持选择性审核,尽管其校准并不一致;Laya的置信度在错误排序方面用处较小。这些结果表明,SOM驱动的审核具有广阔前景。

英文摘要

Online moderation systems must apply changing platform policies, community rules, and prior decisions while producing decisions that can be audited and routed to human review. We evaluate whether System One Models, which accept natural-language context but return typed choices, probabilities, or scores, can support this setting. Across five moderation benchmarks, Jev is competitive with specialized reference systems, matching or exceeding them in several policy-grounded and harmful-content settings. We then use controlled information conditions to separate written rules, retrieved precedents, and restrictions on the candidate answer space. Jev generally benefits from retrieved precedents, improving exact policy selection and harmful-content discrimination when the answer space is held fixed. Laya is less consistent: retrieval often shifts its positive prediction rate or no-violation rate without improving discrimination. Jev confidence can support selective review in several settings, although it is not consistently calibrated; Laya confidence is less useful for ranking errors. These results show a promising future for SOM-powered content moderation.

发表机构

  • Institute of Information Science and Technologies (ISTI)(信息技术科学研究所)
  • National Research Council (CNR)(国家研究委员会)

机构由 AI 辅助整理,请以论文原文为准。

↑