预训练低秩张量分解用于多维图像恢复
Pre-Trained Low-Rank Tensor Decomposition for Multi-Dimensional Image Recovery
查看机构详情
- University of Electronic Science and Technology of China(电子科技大学)
- Pazhou Laboratory (Huangpu)(琶洲实验室(黄埔))
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对张量分解忽略图像间共同结构的问题,提出预训练低秩张量分解框架,融合预训练大视觉模型,在更高恢复保真度下实现更少参数和更低碳足迹,实验性能优于现有方法。
中文摘要 AI 辅助
近年来,张量分解在多维图像表示中十分流行,这类方法从头学习每个图像的实例特定结构。然而,张量分解忽略了不同图像之间的共同结构,导致语义建模能力有限、计算成本高且可学习参数数量庞大。为应对这一挑战,我们提出了首个预训练低秩张量分解(PLTD)框架,该框架将预训练的大视觉模型有机地整合到经典张量分解框架中。与浅层且未经训练的深度张量分解相比,所提出的PLTD在更高的恢复保真度、更少的可学习参数和更小的碳足迹之间实现了前所未有的平衡。具体而言,PLTD将目标张量分解为一个潜在张量和一个可学习变换,该变换将潜在张量映射回原始数据域。潜在张量由两个不可或缺且互补的部分组成,即一个固定的预训练潜在张量和一个可学习的低秩潜在张量。固定的预训练潜在张量从预训练的大视觉模型(即DINOv3)中蒸馏而来,用于捕获目标张量的共同结构,而可学习的低秩潜在张量则表征目标张量的实例特定结构。为检验PLTD的潜力,我们开发了相应的多维图像恢复模型,并从理论上论证了该框架的优势。此外,我们讨论了PLTD与经典张量分解框架之间的联系。在多维图像恢复上的大量实验表明,与最先进的方法相比,PLTD consistently实现了更优越的性能。
英文摘要
Recently, tensor decompositions are prevalent for multi-dimensional image representation, which learn the instance-specific structure of each image from scratch. However, tensor decompositions neglect the common structure across different images, leading to limited semantic modeling capability, high computational cost, and a large number of learnable parameters. To address this challenge, we suggest the first pre-trained low-rank tensor decomposition (PLTD) framework, which organically integrates the pre-trained large vision model into the classical tensor decomposition framework. Beyond the shallow and untrained deep tensor decomposition, the suggested PLTD achieves an unprecedented balance among higher recovery fidelity, fewer learnable parameters, and smaller carbon footprint. Specifically, PLTD factorizes the target tensor into a latent tensor and a learnable transform that maps the latent tensor back to the original data domain. The latent tensor consists of two indispensable and complementary terms, i.e., a fixed pre-trained latent tensor and a learnable low-rank latent tensor. The fixed pre-trained latent tensor is distilled from a pre-trained large vision model (i.e., DINOv3) to capture the common structure of the target tensor, while the learnable low-rank latent tensor characterizes the instance-specific structure of the target tensor. To examine the potential of PLTD, we develop the corresponding multi-dimensional image recovery model and theoretically justify the advantages of this framework. Additionally, we discuss the connections between PLTD and classical tensor decomposition frameworks. Extensive experiments on multi-dimensional image recovery demonstrate that PLTD consistently achieves superior performance compared with state-of-the-art methods.