无需模型:文本熵率过滤缓解迭代微调崩溃
No Model Required: Text Entropy Rate Filtering Mitigates Iterative Fine-Tuning Collapse
浏览论文内容
中文总结 AI 辅助
提出基于信息论的非参数熵率过滤器,无需模型即可在合成数据迭代微调中缓解模型崩溃,显著提升文本多样性并有效检测崩溃。
中文摘要 AI 辅助
在合成数据上进行迭代微调会导致“模型崩溃”:随着稀有模式逐渐丢失,输出多样性收窄,这一特征最明显地表现为短语级别的重复。现有的缓解方法要么需要模型对数概率、外部预言机,要么需要持续访问真实人类数据。在此,我们开发了一种基于数学信息论的新方法:非参数 Kontoyiannis 熵率估计器 $h_k$,完全通过匹配长度统计从原始文本计算,无需任何模型。我们证明,在全合成、单谱系微调设置中,这实际上是一个在文本多样性指标上更优的训练数据过滤器。在 Llama-3.1-8B 上的六代 QLoRA 崩溃实验中,基于对数概率的过滤(最成熟的需模型访问的基线)在任何指标上均未提供显著的文本多样性收益($p > 0.23$),而 $h_k$ 过滤带来了 $+42\%$ 的唯一三元组、$+30\%$ 的词汇量以及 $-19\%$ 的重复率(所有 $p < 0.001$)。我们验证了 $h_k$ 作为跨领域熵代理($\eta = 0.924$,$R^2 = 0.746$)和崩溃检测器($\ ho = +0.454$,$p < 0.0001$)的有效性,涉及 4 个领域、2 个温度、2 个生成器-评分器模型对以及 1,520 个生成文档。我们的结果表明,信息论方法在缓解崩溃方面是高效的,并为维持多智能体多样性提供了新思路。
英文摘要
Iterative fine-tuning on synthetic data causes \emph{model collapse}: output diversity narrows as rare patterns are progressively lost, a signature most visible as phrase-level repetition. Existing mitigations either require model log-probabilities, an external oracle, or continued access to real human data. Here we develop a new approach grounded in mathematical information theory: the non-parametric Kontoyiannis entropy rate estimator $h_k$, computed entirely from raw text via match-length statistics, with no model of any kind. We show that this is in fact a \emph{superior} training-data filter on text-diversity metrics in a fully-synthetic, single-lineage fine-tuning setting. In a six-generation QLoRA collapse experiment on Llama-3.1-8B, logprob-based filtering (the most established model-access-requiring baseline) provides no significant text-diversity benefit on any metric ($p > 0.23$), whereas $h_k$-filtering yields $+42\%$ unique trigrams, $+30\%$ vocabulary, and $-19\%$ repetition (all $p < 0.001$). We validate $h_k$ as a cross-domain entropy proxy ($β= 0.924$, $R^2 = 0.746$) and collapse detector ($ρ= +0.454$, $p < 0.0001$) across 4~domains, 2~temperatures, 2~generator--scorer model pairs, and 1{,}520 generated documents. Our results demonstrate that information theoretic approaches to collapse mitigation are efficient, and suggest new approaches for maintaining multi-agent diversity.
发表机构
- Adelaide Data Science Centre(阿德莱德数据科学中心)
- School of Mathematical Sciences(数学科学学院)
- Adelaide University(阿德莱德大学)
机构由 AI 辅助整理,请以论文原文为准。