arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当数据工件不是捷径:合成RLVR语料库的因果审计

When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora

Esther Xin

arXiv 2610.00202首次发表:更新:

发表机构

Independent Researcher(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究审计合成RLVR语料库中正确性与来源的纠缠,发现表面统计信号可检测但模型在固定预算下未利用该工件,并发布CPU友好审计协议。

AI 中文摘要

近期的若干流水线通过遮蔽真实语料文本中的一段跨度,并让语言模型围绕该跨度编造看似合理的错误答案来构建RLVR训练数据。因此,正确选项是真正的人类文本;而每个干扰项都是合成的。正确性与来源变得纠缠不清,策略原则上可能学习后者而非前者。我们在GooseReason-0.7M中审计了这种可能性。首先,我们探究这种不对称性是否可见:一个仅读取五个表面统计量(从不涉及语义)的分类器在315,499个选项上达到AUROC 0.562,仅略高于随机水平。然而,聚合结果掩盖了某些细节。代码类别的AUROC为0.416,低于随机水平,人工检查解释了原因:代码干扰项实际上是正确答案的单操作符变异,而非自由编写的替代项,因此两类在构造上几乎相同。检测到信号并不等同于表明模型使用了它,因此我们随后进行了一项干预实验。我们构建了一个释义匹配的对照语料库,保持各臂的训练集规模相同,并在固定预算下训练两个策略。利用差距并未偏向未修改数据臂:其值为0.021,而对照组为0.027。换言之,在我们的预算下,一个可检测的工件未被利用。我们认为,这种分离现象以及特定领域的构造发现,对于任何策划此类语料库的人来说都值得了解,并且我们以主要仅依赖CPU的协议形式发布了该审计。

英文摘要

Several recent pipelines build RLVR training data by masking a span of real corpus text and asking a language model to invent plausible wrong answers around it. The correct option is therefore genuine human prose; every distractor is synthetic. Correctness and provenance become entangled, and a policy could in principle learn the second instead of the first. We audit that possibility in GooseReason-0.7M. First we ask whether the asymmetry is visible at all: a classifier reading only five surface statistics (never the meaning) reaches AUROC 0.562 over 315,499 options, barely above chance. The aggregate hides something, though. Code sits at 0.416, below chance, and manual inspection explains why: code distractors turn out to be single-operator mutations of the gold answer rather than freely written alternatives, so the two classes are nearly identical by construction. Detecting a signal is not the same as showing a model uses it, so we then run an intervention. We build a paraphrase-matched control corpus, hold training-set size identical across arms, and train two policies under one fixed budget. The exploitation gap does not favour the unmodified-data arm: 0.021 against 0.027 for the control. Under our budget, in other words, a detectable artifact went unexploited. We think that dissociation, along with the domain-specific construction finding, is worth knowing for anyone curating corpora of this kind, and we release the audit as a mostly CPU-only protocol.

Comments9 pages, 2 figures,4 tables;Code and data https://github.com/ethxin0011/rlvr_authenticity_audit

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑