arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

In-Token Learning:基于扩散Transformer的高保真图像恢复

In-Token Learning for High-Fidelity Image Restoration via Diffusion Transformers

Xingfu Yi, Xiaoxue Yu

arXiv 2609.33523首次发表:更新:

发表机构

Zhejiang University; Independent Researcher(浙江大学; 独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出In-Token Learning框架,通过条件整流流匹配适配扩散Transformer,融合降质令牌与潜在令牌并采用直接低质量引导,在多个数据集上实现高保真图像恢复,含QHD推理与12K演示。

AI 中文摘要

我们提出了In-Token Learning,一种图像恢复框架,它利用条件整流流匹配来适配预训练的扩散Transformer。干净目标与降质输入配对,监督从高斯噪声到恢复图像的传输过程。空间对齐的降质图像令牌沿通道维度与不断演化的潜在令牌融合,在给定分辨率下保持图像令牌数量不变。直接低质量引导(DLG)通过原生条件路径将冻结的降质图像嵌入与固定任务提示相结合,无需可训练的ControlNet风格分支或图像描述。我们在DIV2K、LSDIR、FFHQ、RealLQ250和RealPhoto60上评估了超分辨率和去噪,并在DIV2K和LSDIR上评估了自动着色。这些任务使用同一框架下分别训练的检查点。结果显示,在所评估的协议下具有竞争力的保真度和感知质量,但在RealLQ250上泛化能力较弱。我们报告了全图像QHD(2560×1440)推理以及《清明上河图》的平铺12K恢复演示。注意力成本仍随分辨率增加而增加。本技术报告保留了Fill2SR背后的早期更广泛研究,该研究随后发展了真实世界超分辨率方向。

英文摘要

We present In-Token Learning, an image restoration framework that adapts a pretrained diffusion transformer using conditional rectified flow matching. Clean targets paired with degraded inputs supervise transport from Gaussian noise to restored images. Spatially aligned degraded-image tokens are fused with evolving latent tokens along the channel dimension, preserving the image-token count at a given resolution. Direct Low-Quality Guidance (DLG) combines frozen degraded-image embeddings with a fixed task prompt through the native conditioning pathway, without a trainable ControlNet-style branch or image captioning. We evaluate super-resolution and denoising on DIV2K, LSDIR, FFHQ, RealLQ250, and RealPhoto60, and automatic colorization on DIV2K and LSDIR. The tasks use separately trained checkpoints under the same framework. Results show competitive fidelity and perceptual quality under the evaluated protocols, with weaker generalization on RealLQ250. We report full-image QHD ($2560{\times}1440$) inference and a tiled $12$K restoration demonstration of Along the River During the Qingming Festival. Attention cost still increases with resolution. This technical report preserves the early broader study underlying Fill2SR, which subsequently developed the real-world super-resolution direction.

CommentsTechnical report, 25 pages, 12 figures. Preserves the earlier broader study underlying Fill2SR (ECCV 2026), including automatic colorization; Fill2SR subsequently developed the real-world super-resolution direction

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑