arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12340cs.LG

环境离散扩散:在合适的时间使用错误的数据实现数据高效学习

Ambient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient Learning

Julian Kleutgens, Mauricio Tec, Claudio Battiloro, Francesca Dominici, Giannis Daras

首次发表
浏览论文内容

中文总结 AI 辅助

RefineMix框架在数据匮乏时,于离散扩散的低噪声阶段用分布外数据训练模型,在五个领域偏移设置中表现优于或相当于领域内微调,蛋白质生成任务中仅197个示例就使优质蛋白生成比例近翻倍。

中文摘要 AI 辅助

我们提出RefineMix,这是一个用于在数据极度匮乏的情况下训练离散扩散模型的框架,数据匮乏是科学应用中的常见约束。RefineMix在选定的扩散时间使用分布外数据来提升泛化能力,同时不会使采样分布产生偏差。尽管该策略已在连续扩散中被探索,但离散扩散面临独特挑战:与高斯噪声不同,掩码会在保留的token中保留领域信息,限制了在高噪声水平下使用相关数据;而在低噪声水平下,领域的有效不相交支撑则成为优势,使模型能从领域内和分布外数据中学习,且不会使采样器产生偏差。我们将这些直觉形式化,并为所提方法提供理论分析。实验在五个领域偏移设置中,RefineMix的表现与领域内微调或数据混合相当或更优。在蛋白质序列生成任务中,仅用197个领域内示例进行微调,与标准微调相比,生成的同时具备新颖性、可折叠性和家族内属性的蛋白质比例几乎翻倍。

英文摘要

We introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications. RefineMix uses out-of-distribution data at selected diffusion times to improve generalization without biasing the sampling distribution. Although this strategy has been explored in continuous diffusion, discrete diffusion presents a distinct challenge: unlike Gaussian noise, masking preserves domain information in surviving tokens, limiting the use of related data at high noise levels. At low noise levels, however, the domains effectively disjoint supports become an advantage, allowing the model to learn from both in-domain and out-of-distribution data without biasing the sampler. We formalize these intuitions and provide a theoretical analysis for the proposed method. Experimentally, across five domain-shift settings, RefineMix matches or outperforms in-domain finetuning and data mixing. For protein sequence generation, finetuning with just 197 in-domain examples nearly doubles the fraction of generated proteins that are simultaneously novel, foldable, and in-family compared to standard finetuning.

发表机构

  • ETH Zürich(苏黎世联邦理工学院)
  • Harvard University(哈佛大学)
  • MIT(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑