arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自重建动力学用于自编码器重建精化

Self-Reconstruction Dynamics for Autoencoder Reconstruction Refinement

Hitoshi Iyatomi

arXiv 2609.32268首次发表:更新:

AI 中文总结

本文提出自重建动力学(SRD)引导的重建精化方法,利用冻结自编码器的重复重建轨迹预测潜在修正,无需测试时优化,在六个数据集上恢复约40%的MSE差距并提升PSNR,证实了样本特异性精化信号的有效性。

AI 中文摘要

标准自编码器(AE)推理使用单次编码器-解码器传递,但在固定解码器下,潜在表示可能并非对每个样本都是最优的。我们提出疑问:训练好的AE能否揭示用于改进其自身重建的信息?将冻结的AE重复应用于其重建结果,会产生瞬时的图像空间和潜在空间轨迹,称为自重建动力学(SRD)。尽管在本文研究的AE中这会降低保真度,但SRD包含用于纠正重建的样本特定信息。我们提出SRD引导的重建精化(SRD-RR),该方法从短SRD预测潜在修正,AE保持冻结,无需逐样本测试时优化。我们还引入了MSE-recov,即相对于经验解码器优化参考的MSE恢复比率。在六个数据集上,SRD-RR在一次转换中恢复了经验可恢复MSE差距的38.6%,两次转换恢复45.3%。一种不直接访问原始图像、使用SRD派生的伪目标训练的两转换变体,实现了40.7%的恢复率和1.74 dB的平均PSNR增益。移除轨迹信息会降低增益,而跨样本轨迹分配导致严重退化,证实了强样本特异性。非线性SRD条件精化始终优于固定和训练的线性潜在修正。在基于预训练DINOv2的表示自编码器(RAE)上,其潜在动力学显著不同,SRD条件再次改进了匹配的无轨迹预测器。然而,像素MSE潜在精化揭示了像素保真度与感知质量之间的强烈不匹配,而SRD派生的伪目标缓解了这种退化。总体而言,SRD是一种有用的样本特定精化信号,而目标函数决定了它如何转化为像素和感知质量。

英文摘要

Standard autoencoder (AE) inference uses a single encoder-decoder pass, though the latent may not be optimal for each sample under a fixed decoder. We ask whether a trained AE can reveal information for improving its own reconstruction. Repeated application of a frozen AE to its reconstruction produces transient image- and latent-space trajectories, termed Self-Reconstruction Dynamics (SRD). Although this degrades fidelity in the AEs studied here, SRD contains sample-specific information for correcting the reconstruction. We propose SRD-guided Reconstruction Refinement (SRD-RR), which predicts a latent correction from a short SRD with the AE frozen and no per-sample test-time optimization. We also introduce MSE-recov, an MSE recovery ratio relative to an empirical decoder-optimized reference. Across six datasets, SRD-RR recovers 38.6% of the empirically recoverable MSE gap with one transition and 45.3% with two. A two-transition variant trained without direct access to original images, using an SRD-derived pseudo-target, achieves 40.7% recovery and a 1.74 dB average PSNR gain. Removing trajectory information reduces the gain, while cross-sample trajectory assignment causes severe degradation, confirming strong sample specificity. Nonlinear SRD-conditioned refinement consistently outperforms fixed and trained linear latent correction. On a pretrained DINOv2-based representation autoencoder (RAE) with substantially different latent dynamics, SRD conditioning again improves a matched trajectory-free predictor. However, pixel-MSE latent refinement reveals a strong mismatch between pixel fidelity and perceptual quality, while the SRD-derived pseudo-target mitigates this degradation. Overall, SRD is a useful sample-specific refinement signal, while the objective determines how it translates into pixel and perceptual quality.

Comments23 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑