TAKE:用于文本数据集蒸馏的轨迹感知知识估计
TAKE: Trajectory-Aware Knowledge Estimation for Text Dataset Distillation
浏览论文内容
中文总结 AI 辅助
针对大规模文本语料库瓶颈问题,提出文本数据集蒸馏框架TAKE,通过影响函数及轨迹感知知识估计卷积样本知识分数作权重,在极端压缩下评估,可保持任务保真度实现数据效率,对相关研究有理论意义且开源代码。
中文摘要 AI 辅助
大规模文本语料库已成为现代自然语言处理中的一个瓶颈,不仅在存储方面,还在训练、微调及持续学习的累积成本上。我们提出一种文本数据集蒸馏框架,能将语料库缩减至原大小的0.1%,同时保持下游任务的保真度。我们通过影响函数来进行蒸馏,它能量化每个样本对下游目标的贡献。我们引入轨迹感知知识估计(TAKE),将基于知识的影响沿训练轨迹卷积为单个样本知识分数,捕获信息丰富的样本。这些分数作为离散最优传输目标中的样本权重,指导从合成生成的候选池中选择原型。我们在极端压缩(0.1%或每个类别20个样本)下对文本分类和自然语言推理任务的下游准确性评估TAKE,表明在不牺牲任务保真度的情况下可实现数据效率。该方法有理论基础,对核心集构建和以数据为中心的人工智能有更广泛的意义。我们在这个https网址发布了源代码。
英文摘要
Large-scale text corpora have become a quiet bottleneck in modern NLP, not just in storage, but in the accumulated cost of training, fine-tuning, and continual learning. We propose a text dataset distillation framework that reduces corpora to as little as 0.1% of their original size while preserving downstream task fidelity. We approach distillation through the lens of influence functions, which quantify each sample's contribution to the downstream objective, a natural and principled basis for selection. We introduce Trajectory-Aware Knowledge Estimation (TAKE), which convolves the knowledge-based influence along the training trajectory into a single per-sample knowledge score, capturing informative samples. These scores serve as sample weights within a discrete Optimal Transport objective, guiding prototype selection from a synthetically generated candidate pool. We evaluate TAKE on downstream accuracy across text classification and natural language inference tasks at extreme compression (0.1% or 20 samples/class), showing that data efficiency is achievable without sacrificing task fidelity. The approach is theoretically grounded, with broader implications for coreset construction and data-centric AI. We release our source code at https://github.com/votrinhan88/take.