arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24782cs.CV

基于自编码器潜在表示的图像差异量化

Image Difference Quantification Using Autoencoder-Based Latent Representations

Manish Sharma, Timothy Yim, Clifton Forlines

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出基于卷积自编码器的框架,在潜在空间用余弦相似度量化图像差异,经猫狗等数据集验证,其对感知相关图像失真敏感,可作为传统像素级相似度指标的替代方案,具多领域应用潜力。

中文摘要 AI 辅助

传统图像相似度度量指标,如均方误差(MSE)、峰值信噪比(PSNR)和结构相似性指数测量(SSIM),均依赖像素级比较,常无法捕捉图像间具有感知意义的差异。相比之下,深度神经网络学习得到的潜在表示编码了与人类视觉感知更契合的高级语义信息。本文提出一种基于卷积自编码器的框架,用于在潜在空间中利用余弦相似度量化图像差异。学习得到的紧凑嵌入能够在光照、姿态和背景变化下,对视觉上不同的图像实现鲁棒区分。对猫狗图像及额外跨域数据集的广泛评估表明,潜在空间中存在清晰的类别聚类和强类间可分性,98.4%的猫狗图像对的相似度得分低于0.5。使用TID2013数据集的进一步验证显示,潜在空间距离与人类平均意见得分(MOS)呈正相关,证明其对感知相关图像失真具有敏感性。所提方法为传统基于像素的相似度度量提供了计算高效且语义合理的替代方案,在基于内容的检索、感知质量评估和语义相似度分析等领域具有潜在应用价值。

英文摘要

Traditional image similarity metrics such as Mean Squared Error (MSE), Peak Signal-to-Noise Ratio (PSNR), and the Structural Similarity Index Measure (SSIM) rely on pixel-level comparisons and often fail to capture perceptually meaningful differences between images. In contrast, latent representations learned by deep neural networks encode high-level semantic information that is more closely aligned with human visual perception. This paper proposes a convolutional autoencoder-based framework for quantifying image differences using cosine similarity in latent space. The learned compact embeddings enable robust differentiation between visually distinct images under variations in illumination, pose, and background. Extensive evaluation on dog-cat images and additional cross-domain datasets demonstrates clear class-wise clustering and strong inter-class separability in the latent space, with 98.4% of dog-cat image pairs exhibiting similarity scores below 0.5. Further validation using the TID2013 dataset shows that latent-space distance correlates positively with human Mean Opinion Scores (MOS), demonstrating sensitivity to perceptually relevant image distortions. The proposed approach provides a computationally efficient and semantically grounded alternative to conventional pixel-based similarity metrics, with potential applications in content-based retrieval, perceptual quality assessment, and semantic similarity analysis.

↑