arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当保留约束冲突时:缓解表格数据中的遗忘-保留干扰

When Retain Constraints Conflict: Mitigating Forget-Retain Interference in Tabular Data

Zijie Liu, Jinhao Duan, Bingqi Shang, Xinming An, Sijia Liu, Tianlong Chen

arXiv 2609.06786首次发表:更新:

发表机构

University of North Carolina at Chapel Hill; Michigan State University(北卡罗来纳大学教堂山分校; 密歇根州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对表格数据遗忘中模式导致的遗忘-保留冲突,提出冲突感知遗忘(CAU)方法,通过放宽冲突保留行约束,更接近重训练基准并保持效用。

AI 中文摘要

机器遗忘旨在移除指定训练数据的影响,同时保持模型效用,但其在表格数据上的行为仍未得到充分探索。这一空白之所以重要,是因为表格预测广泛应用于高风险领域,并日益通过记录序列化和模式感知提示适配到语言模型。我们识别出一个关键挑战,它将表格遗忘与自由文本或其他模态中的遗忘区分开来:模式诱导的遗忘-保留重叠。在序列化的表格数据中,记录共享固定的列名/值槽位、相似的属性范围和共同的输出空间。因此,一个遗忘行可能有邻近的保留行依赖相同的高信号属性,导致保留保持与遗忘所需的更新相对立。受此失败模式的启发,我们提出了冲突感知遗忘(CAU),一种模式感知的方法,通过放宽与遗忘集最冲突的保留行上的保留约束来减少遗忘-保留干扰。在临床和非医疗表格任务上的样本级和特征级遗忘实验中,CAU更接近重训练基准,同时保持预测效用和保留区域行为。我们的结果表明,可靠的表格大语言模型遗忘不仅依赖于遗忘目标,还依赖于保留约束的构建方式。

英文摘要

Machine unlearning aims to remove the influence of designated training data while preserving model utility, but its behavior on tabular data remains underexplored. This gap is important because tabular prediction is widely used in high-stakes domains and is increasingly adapted to language models through record serialization and schema-aware prompting. We identify a key challenge that distinguishes tabular unlearning from unlearning in free-form text or other modalities: schema-induced forget-retain overlap. In serialized tabular data, records share fixed column-name/value slots, similar attribute ranges, and common output spaces. Consequently, a forget row may have nearby retain rows that rely on the same high-signal attributes, causing retain preservation to oppose the update required for forgetting. Motivated by this failure mode, we propose Conflict-Aware Unlearning (CAU), a schema-aware approach that reduces forget-retain interference by relaxing preservation constraints on retained rows that most conflict with the forget set. Across sample-level and feature-level unlearning on clinical and non-medical tabular tasks, CAU more closely matches a retraining oracle while maintaining predictive utility and retain-region behavior. Our results show that reliable tabular LLM unlearning depends not only on the forgetting objective, but also on how retain constraints are constructed.

CommentsEMNLP 2026 Finding Paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑