arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2302.04246cs.LGcs.CV

基于变分自编码器的捷径检测

Shortcut Detection with Variational Autoencoders

Nicolas M. Müller, Simon Roschmann, Shahbaz Khan, Philip Sperl, Konstantin Böttinger

首次发表 更新
浏览论文内容

中文总结 AI 辅助

针对机器学习模型易依赖数据虚假相关性(捷径)的问题,提出基于变分自编码器(VAE)的捷径检测方法,利用VAE潜空间特征解耦发现并半自动评估特征-目标相关性,在多个真实数据集上验证有效并发现新捷径。

中文摘要 AI 辅助

对于机器学习(ML)的实际应用而言,模型基于泛化性良好的特征而非数据中的虚假相关性做出预测至关重要。这类虚假相关性也被称为捷径(shortcut),其识别是一项具有挑战性的问题,迄今鲜有研究涉及。在本工作中,我们提出了一种利用变分自编码器(VAE)检测图像和音频数据集中捷径的新方法。VAE潜空间中的特征解耦使我们能够发现数据集中的特征-目标相关性,并半自动评估它们是否为ML捷径。我们在多个真实世界数据集上验证了该方法的适用性,并识别出了此前未被发现的捷径。

英文摘要

For real-world applications of machine learning (ML), it is essential that models make predictions based on well-generalizing features rather than spurious correlations in the data. The identification of such spurious correlations, also known as shortcuts, is a challenging problem and has so far been scarcely addressed. In this work, we present a novel approach to detect shortcuts in image and audio datasets by leveraging variational autoencoders (VAEs). The disentanglement of features in the latent space of VAEs allows us to discover feature-target correlations in datasets and semi-automatically evaluate them for ML shortcuts. We demonstrate the applicability of our method on several real-world datasets and identify shortcuts that have not been discovered before.

发表机构

  • Fraunhofer AISEC(弗劳恩霍夫人工智能安全研究所)
  • Technical University of Munich(慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑