arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAGE:基于视觉基础表征的显著因子发现与生成

SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations

Shuang Liang, Lejun Liao, Shiyuan Zhang, Max C. Zhang, Xiaolong Luo, Han Wang, Stefano Anzellotti, Yuan Yuan

arXiv 2609.39635首次发表:更新:

发表机构

HKU; Boston College; Harvard University; University of Virginia(香港大学; 波士顿学院; 哈佛大学; 弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SAGE在冻结自编码器的空间潜变量中学习显著与常见因子,以显著表征条件化扩散Transformer,实现高保真重建、无监督子类型发现和条件生成,在Digits-ImageNet和FFHQ眼镜上优于基线。

AI 中文摘要

给定一个目标数据集(例如戴眼镜的人脸)和一个背景数据集(例如不戴眼镜的人脸),对比分析将目标特有的显著因子与两者共有的常见内容区分开来。我们旨在获得能够捕捉每张图像中目标特定细节(如眼镜的形状、颜色和位置)的显著表征,以便在没有子类型标签的情况下揭示子类型,并指导生成新发现的子类型的新样本,即使该子类型没有名称或文本描述。我们提出了SAGE,它直接在冻结的表征自编码器的高维空间潜变量中学习这两种因子,并将扩散Transformer以参考图像的已学习显著表征为条件。在Digits-ImageNet和FFHQ眼镜数据集上,SAGE结合了高保真重建(rFID低于2)与无监督子类型发现,比基线更好地恢复了数字(探针准确率0.950对比最高0.281),并揭示了眼镜类型、更精细的太阳镜样式和错误标记的图像;显著条件生成将Digits-ImageNet的子类型准确率从非分解潜变量的27.7%提高到90.5%,并在两个数据集上提高了多样性。在视网膜OCT上,SAGE的显著空间仅使用正常/疾病标签就能区分三种疾病。

英文摘要

Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such as the shape, color, and position of the glasses, so that they reveal subtypes without subtype labels and guide the generation of new examples of a discovered subtype, even one with no name or text description. We introduce SAGE, which learns both factors directly in the high-dimensional spatial latent of a frozen representation autoencoder and conditions a diffusion transformer on the learned salient representation of a reference image. On Digits-ImageNet and FFHQ eyeglasses, SAGE combines high-fidelity \textit{reconstruction} (rFID below $2$) with unsupervised \textit{subtype discovery}, recovering the digits better than baselines (probe accuracy $0.950$ vs.\ at most $0.281$) and revealing eyewear types, finer sunglasses styles, and mislabeled images; salient-conditioned \textit{generation} raises Digits-ImageNet subtype accuracy over the unfactorized latent ($90.5\%$ vs.\ $27.7\%$) and diversity on both datasets. On retinal OCT, SAGE's salient space separates three diseases using only normal/disease labels.

Comments28 pages, 18 figures, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑