arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OriginBlame:人工智能训练数据集的记录级和令牌级数据溯源

OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets

Haolin Xue

arXiv 2607.13037首次发表:更新:

发表机构

HuggingFace; Datatrove; The Linux Foundation; Databricks; Weights & Biases(拥抱脸; 数据宝库; Linux基金会; Databricks公司; 权重与偏差公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对模型训练者在数据删除时无法精准定位记录所属作者的问题,提出OriginBlame系统,通过记录级和令牌级溯源及确定性查询解决,评估显示能消除过度删除并提升遗忘效果,增加少量吞吐量开销。

AI 中文摘要

当数据贡献者要求删除数据时,模型训练者面临一个实际问题:遗忘算法需要一个遗忘集,但没有工具能定位哪些训练记录属于给定作者。现有溯源系统在文件或数据集级别运行,导致过度删除。我们提出了ob,一个记录级和令牌级数据溯源系统,它通过数据处理管道传播作者身份,并通过确定性查询将撤销请求解析为精确的遗忘集。对219,555个维基百科页面的评估表明,记录级溯源消除了数据集级的过度删除,集成增加了吞吐量开销,基于溯源的遗忘集在1.7B模型上比随机基线提高了42%的遗忘效果。

英文摘要

When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can locate which training records belong to a given author. Existing provenance systems operate at file or dataset level, forcing catastrophic over-deletion. We present ob, a record- and token-level data provenance system that propagates author identity through data processing pipelines and resolves revocation requests into precise forget sets via deterministic queries. Evaluation on 219,555 Wikipedia pages demonstrates that record-level provenance eliminates dataset-level over-deletion (from 101x to 1.3x), while integration adds 1.3-4.0% throughput overhead (HuggingFace) and 2.1-19.0% (Datatrove) on wiki data. On a 1.7B model, provenance-based forget sets consistently reduce the collateral damage of machine unlearning (retain perplexity) relative to same-size random baselines across all evaluated authors, with membership-inference tests indicating they select genuinely memorized content.

Comments13 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑