arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

方向影响函数:估计约束学习中的训练数据影响

Directional Influence Function: Estimating Training Data Influence in Constrained Learning

Xin Wang, R. Tyrrell Rockafellar, Xuegang, Ban

arXiv 2607.23388首次发表:更新:

发表机构

University of Washington(华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究约束学习中训练数据对模型解的影响,提出方向影响函数(DIF),将约束学习最优性条件表述为变分不等式,在约束线性回归和公平性约束的卷积神经网络上验证,结果表明DIF是约束学习中数据归因的有效可靠工具。

AI 中文摘要

随着约束学习越来越普遍,模型在明确的可行性要求下进行训练,以确保公平性、安全性、鲁棒性、正则化以及物理或逻辑约束。理解训练样本如何影响模型解(如学习到的参数)对于可解释性和鲁棒性至关重要。经典影响函数(IF)通过局部敏感性分析估计样本贡献,但在约束设置中不可靠,因为数据扰动会重塑目标和可行区域,导致违反可行性的估计。为此,我们提出方向影响函数(DIF),一种将这些约束明确纳入影响估计的新估计器。DIF将约束学习的最优性条件表述为变分不等式(VI),并分析扰动训练数据如何影响该VI。我们在约束线性回归上验证了DIF,表明它能恢复留一法再训练结果,而IF和基于惩罚的IF存在显著偏差。我们进一步将DIF应用于公平性约束的卷积神经网络,DIF能准确预测数据删除后的测试损失变化,并与实际再训练紧密对齐。我们的结果确立了DIF作为约束学习中数据归因的有效且可靠工具。

英文摘要

As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustness, regulariza- tion, and physics or logic constraints. Understanding how training samples in- fluence the model solution (e.g., learned parameters) is crucial for interpretability and robustness. The classical influence function (IF) estimates sample contribu- tions via local sensitivity analysis, measuring how the solution changes when a specific training sample is perturbed or removed. However, IF becomes unreli- able in constrained settings: data perturbations can reshape both the objective and the feasible region, leading to estimates that violate feasibility. In response, we propose the Directional Influence Function (DIF), a novel estimator that explicitly incorporates these constraints into influence estimation. DIF formulates the opti- mality conditions of constrained learning as a variational inequality (VI) and ana- lyzes how perturbing training data affects this VI. We validate DIF on constrained linear regression and demonstrate that it recovers leave-one-out retraining results, whereas IF and penalty-based IF exhibit significant bias. We further apply DIF to fairness-constrained CNNs, where DIF accurately predicts test loss changes under data removal and aligns closely with actual retraining. Our results establish DIF as an efficient and reliable tool for data attribution in constrained learning.

CommentsNeed revision

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑