arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于记忆、提取与版权的误解:打地鼠游戏

Playing Whack-a-Mole with misconceptions about memorization, extraction, and copyright

A. Feder Cooper

arXiv 2609.09320首次发表:更新:

AI 中文总结

本文指出《对齐打地鼠》论文中微调记忆测量方法无效,存在假阳性风险,其版权提取主张缺乏证据支持。

AI 中文摘要

经过仔细审查,我确信《对齐打地鼠》中的头条微调记忆结果采用了无效的测量程序。这些头条结果所依赖的书籍记忆覆盖率指标,其计数的序列匹配长度远短于领域标准所认可的作为记忆有效证据的长度,并且用于诱发记忆的提示程序存在在提示中泄露被“提取”文本的风险。该论文未包含必要的阴性对照实验,以判断结果在多大程度上因假阳性而膨胀:当生成内容与训练数据之间的匹配可能源于其他因素时,却声称提取成功(从而证明训练数据的记忆)。鉴于这些有效性问题,论文声称微调能让用户以可替代原版的形式提取受版权保护书籍的实质性部分,这一结论并未得到所报告结果的支持。未报告实验成本(威胁模型的重要组成部分)进一步削弱了版权主张。我撰写此说明是因为在过去一个月里,(潜在)原告已联系我询问这篇论文。他们希望引用这项工作作为正在进行的和未来可能的版权诉讼中支持其主张的有效证据。

英文摘要

After careful review, I'm confident the headline fine-tuning memorization results in Alignment Whack-a-Mole use an invalid measurement procedure. The book memorization coverage metric these headline results depend on counts sequence matches far shorter than what field standards consider valid evidence of memorization, and the prompting procedure used to elicit memorization runs the risk of leaking the text being "extracted" in the prompt. The paper doesn't include the negative-control experiments needed to see how much the results are inflated by false positives: claiming extraction success (and therefore memorization of training data) when matches between generations and training data may be due to other factors. Given these validity issues, the paper's claims that fine-tuning lets users extract substantial portions of copyrighted books, in a form that could substitute for the originals, aren't supported by the reported results. The failure to report the experiments' cost (an important component of the threat model) further compromises the copyright claims. I'm writing this note because, in the last month, (prospective) plaintiffs have reached out to me to ask about this paper. They're looking to cite this work as valid evidence in support of claims in ongoing and potential future copyright litigation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑