arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

来源识别不是适应性测试:衡量合成数据归因的极限

Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution

Joss Armstrong

arXiv 2610.00417首次发表:更新:

发表机构

Ericsson(爱立信)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过金融风险文本实验证明,合成数据来源归因在改写后准确性大幅下降,且来源信息与数据训练价值无关,二者是独立问题,现有方法均不足以支撑递归训练。

AI 中文摘要

在模型生成的数据上反复训练可能会降低后续模型的性能。一种可能的应对措施是在决定重用哪些生成示例时使用来源信息。我们测试了这种来源信息恢复的可靠性,以及它是否有助于识别更好的训练数据。使用金融风险文本,我们首先识别生成段落的来源,然后在改写后重复测试。生成器归因在原始段落上的准确率为98.7%,但在释义后降至53.1%,在风格改写后降至29.0%。生成与人类检测在测试的人类比较集上仍接近完美。然后,我们在三轮生成和再训练中比较了两种选择生成示例的方法。一种使用来源信息,另一种使用来自独立参考模型的评分。这两种规则选择了不同的示例,但计划中的比较并未检测到由此产生的模型退化中存在稳定的差异。结果表明,识别数据来源和识别哪些数据对训练有用是两个独立的问题。因此,该实验将来源身份、面向标准的筛选和递归训练结果区分开来:既未证明来源评分,也未证明所测试的面向标准的代理足以支持未来的递归行为。

英文摘要

Repeated training on model-generated data can degrade later models. One possible response is to use provenance when deciding which generated examples to reuse. We test both how reliably that provenance can be recovered and whether it helps identify better training data. Using financial-risk text, we first identify the source of generated passages and then repeat the test after rewriting them. Generator attribution is 98.7% accurate on the original passages but falls to 53.1% after paraphrasing and 29.0% after style rewriting. Generated-versus-human detection remains close to perfect against the tested human comparison set. We then compare two ways of selecting generated examples over three rounds of generation and retraining. One uses source information. The other uses a score from a separate reference model. The two rules select different examples, but the planned comparison does not detect a stable difference in the degradation of the resulting models. The results show that identifying where data came from and identifying which data are useful for training are separate problems. The experiment therefore separates source identity, criterion-facing selection, and recursive training outcome: neither the provenance score nor the tested criterion-facing proxy is established as sufficient for future recursive behaviour.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑