arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22522cs.LG

从有害到有益:基于动态影响力的评估与编辑

From Detrimental to Beneficial: Dynamic Influence-based Valuation and Editing

Adrian Nyakairu, Hongfu Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出DIVE框架,通过在优化层面反转有害样本梯度方向,将有害数据转化为有益贡献,可提升分类性能、最大化数据效率并稳定优化,还能泛化到大语言模型微调。

中文摘要 AI 辅助

数据评估是数据中心学习的基石,现有研究主要聚焦于设计算法将训练样本分类为对学习任务有益或有害,但利用这些评估结果进行后续数据干预的研究仍未得到充分探索;传统方法通常丢弃或降低有害样本的权重,从而未能充分利用可用数据资源。本文提出了一种新颖高效的框架——基于动态影响力的评估与编辑(DIVE),该框架在批次级别动态估计样本价值,并将有害数据转化为有益贡献。DIVE不改变原始数据,而是在优化层面操作,通过在训练期间策略性地反转有害样本的梯度方向,确保与标准学习流程无缝集成且开销极小。大量实证评估表明,DIVE可持续提升分类性能、最大化数据效率、稳定优化过程,并能有效泛化到大型语言模型微调任务中。

英文摘要

Data valuation is a cornerstone of data-centric learning, where prior efforts primarily focus on designing algorithms to classify training samples as either beneficial or detrimental for the learning task. However, leveraging these valuation estimates for subsequent data intervention remains underexplored; conventional approaches typically discard or downweight harmful samples, thereby underutilizing available data resources. In this paper, we present Dynamic Influence-based Valuation and Editing (DIVE), a novel and efficient framework that dynamically estimates sample values at the batch level and transforms detrimental data into beneficial contributions. Rather than altering the raw data, DIVE operates at the optimization level by strategically reversing the gradient directions of harmful samples during training, ensuring seamless integration with standard learning procedures with minimal overhead. Extensive empirical evaluations demonstrate that DIVE consistently improves classification performance, maximizes data efficiency, stabilizes optimization, and effectively generalizes to large language model fine-tuning.

发表机构

  • Brandeis University(布兰迪斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑