arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.03788cs.LG

用于少步离散扩散的张量训练联合建模

Tensor-Train Joint Modeling for Few-Step Discrete Diffusion

  • KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

Byoungkwon Kim, Minhyuk Sung

AI总结:

研究离散扩散少步生成潜力受限问题,用张量分解进行联合分布建模,支持多种分解方式,提出迭代边际推理程序,通过微调预训练模型大幅提升少步生成效果。

AI中文摘要:

离散扩散有望比自回归模型更快生成顺序离散数据,但由于结构限制其少步生成潜力未被充分挖掘。本文提出首个通过张量分解进行离散扩散显式联合分布建模的框架,支持多种分解方式,识别出TTD对附近令牌依赖的结构偏差,提出迭代边际推理程序,通过微调预训练模型提升少步生成效果。

英文摘要:

Discrete diffusion promises orders-of-magnitude faster generation than autoregressive (AR) models for sequential discrete data, yet its full potential of few-step generation has remained out of reach due to a fundamental structural limitation. The conditional-independence assumption underlying current discrete diffusion models introduces a systematic parallelization bias that compounds with the number of tokens unmasked per step, becoming severe in the few-step regime that fast generation requires. We address this with the first framework for explicit joint distribution modeling in discrete diffusion via tensor decomposition, which represents the conditional clean distribution as a low-rank tensor with controllable expressivity. The framework supports both Canonical Polyadic (CPD) and Tensor-Train (TTD) decompositions, and we identify a structural bias of TTD toward dependencies between nearby tokens, formalized through Oseledets' theorem relating TT-rank to unfolding-matrix rank, which is well-suited to sequential data such as natural language and line notations for molecular data. To enable efficient generation, we present an iterative marginal inference procedure with specialization for predetermined position schedules. Our framework integrates into pretrained MDMs through lightweight fine-tuning, yielding substantial improvements in few-step generation at a fraction of the cost of training from scratch. Code available at https://github.com/ssamt/tensor-train.

↑