arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SlopBench:我们能在多大程度上根据“AI 废料”对语言模型进行排序?一个多领域的重复性 AI 写作基准

SlopBench: How Well Can We Rank Language Models by Slop? A Multi-Domain Benchmark of Repetitive AI Writing

Dhruv Roongta, Harsha Gaddipati, Anh Tuan Huynh

arXiv 2609.33905首次发表:更新:

发表机构

Slashy(Slashy)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出 SlopBench 基准,通过四种表面行为对 18 个模型在 112 个任务上的 19,928 个输出进行评分,发现 Kimi K2.6 得分最低、Mistral Large 最高,但排序不稳定,故分别报告各行为并公开数据与代码。

AI 中文摘要

SlopBench 探究哪些模型会生成读者称之为“AI 废料”的僵硬、重复的散文,这是一个在检测器将文本归类为机器撰写后仍未解答的问题。我们在电子邮件、社交媒体帖子、论文和工作场所聊天中,对 18 个模型在 112 个手写任务上进行了评估,每个模型在每个任务上最多采样十次,共计 19,928 个输出。SlopBench 对读者可手动检查的四种表面行为进行评分:长度是否符合每个任务指定的词数区间、同一任务中模型自身样本的开头重复情况、段落节奏以及针对 ChatGPT 之前人类语料库的固定词汇结构。在一种固定权重下,Kimi K2.6 得分最低,为 21.1,而 Mistral Large 得分最高,为 40.6。在 500 次随机重新加权中,Kimi 在 58% 的抽样中得分最低,而 Mistral 在 97% 的抽样中得分最高。没有任何一次抽样能保持这 18 个模型的完整排序,而情景自助法仅留下其中一个排名是明确的。我们对该中间排序进行了三项进一步检查:众包竞技场、AI 检测器和词汇多样性。这些检查均未确认该排序。因此,我们分别报告这四种行为,并将综合得分视为众多权重中的一种,同时我们发布了提示词、输出、参考统计数据和评分代码。

英文摘要

SlopBench asks which models produce the stiff, repetitive prose readers call AI slop, a question detectors leave open once they have classified a text as machine-written. We evaluated eighteen models on 112 hand-written tasks in email, social posts, essays, and workplace chat, sampling each model on each task up to ten times, for 19,928 outputs in all. SlopBench scores four surface behaviors a reader can check by hand: length against the word band each task specifies, opener repetition across a model's own samples of one task, and paragraph rhythm and fixed lexical constructions against pre-ChatGPT human corpora. Under one fixed weighting, Kimi K2.6 scores lowest at 21.1 and Mistral Large highest at 40.6. Across 500 random reweightings Kimi has the lowest score in 58 percent of draws and Mistral the highest in 97 percent. No draw preserves the full order of the eighteen, and a scenario bootstrap leaves exactly one of those ranks unambiguous. We ran three further checks on that middle order: a crowd arena, an AI detector, and lexical diversity. None of them confirmed the order. We therefore report the four behaviors separately and treat the composite as one weighting among many, and we release the prompts, outputs, reference statistics, and scoring code.

Comments12 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑