arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向视觉基础模型的快速解耦反事实生成

Towards Fast and Disentangled Counterfactuals for Visual Foundation Models

Sidney Bender, Benedikt Kunz, Ahmed Zeid, Shinichi Nakajima, Klaus-Robert Müller, Marco Morik

arXiv 2610.00895首次发表:更新:

发表机构

Technische Universität Berlin; Berlin Institute for the Foundations of Learning and Data (BIFOLD)(柏林工业大学; 柏林学习与数据基础研究所(BIFOLD))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出解耦扩散自编码器(DiDAE),通过闭式编辑解耦字典方向生成反事实,无需梯度,速度提升2000倍,能修复下游分类器并识别虚假相关性。

AI 中文摘要

基础模型仍然容易受到虚假相关性和“聪明汉斯”策略的影响。可解释机器学习可以在没有元数据的情况下为分类器找到并移除这些策略。对于基础模型,目前还没有这样的选项。我们提出了解耦扩散自编码器(DiDAE)。DiDAE将冻结的基础模型包裹在条件扩散解码器中。反事实是对解耦字典中一个方向的闭式编辑,随后进行解码。字典可以是有监督的(Procrustes)或无监督的(奇异值分解、稀疏自编码器)。由于不需要梯度,DiDAE比现有技术快最多2000倍。我们在六个数据集上进行了评估,其中两个为合成数据集,四个为真实世界数据集。在基于需求驱动的基准测试中,在其中的三个数据集上,其反事实与现有技术相当或更优,并且它们通过反事实知识蒸馏(CFKD)修复下游分类器,在此过程中超越了基于元数据的修正。同一机制可以对预训练字典与训练好的分类器进行排序。它返回分类器实际读取的少数方向,每个方向都通过翻转决策的反事实进行因果验证,并沿着教师标记为虚假的方向修复分类器。该工作流在我们随论文发布的开源Peal库中即插即用。有了公共字典和预训练解码器,剩下的只是分类器的廉价线性蒸馏及其自身的微调。

英文摘要

Foundation models remain vulnerable to spurious correlations and ``Clever Hans'' strategies. Explainable machine learning can find and remove such strategies for classifiers without metadata. For foundation models, no such option exists yet. We propose Disentangled Diffusion Autoencoders (DiDAE). DiDAE wraps a frozen foundation model in a conditional diffusion decoder. A counterfactual is one closed-form edit along a direction of a disentangled dictionary, followed by decoding. The dictionary can be supervised (Procrustes) or unsupervised (Singular Value Decomposition, Sparse Autoencoders). No gradients are needed, so DiDAE is up to 2000 times faster than the state of the art. We evaluate on six datasets, two synthetic and four real-world. In a desiderata-driven benchmark on three of them, its counterfactuals are on par with or better than the state of the art, and they repair downstream classifiers through Counterfactual Knowledge Distillation (CFKD), where they beat metadata-based correction. The same machinery can rank a pretrained dictionary against a trained classifier. It returns the few directions the classifier actually reads, each causally verified by a counterfactual that flips the decision, and repairs the classifier along those a teacher marks spurious. The workflow is plug-and-play in our open-source Peal library we publish alongside the paper. With a public dictionary and a pretrained decoder, all that remains is a cheap linear distillation of the classifier and its own fine-tuning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑