arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FailSAE:基于稀疏自编码器的视觉语言模型可解释故障预测研究

FailSAE: Towards Interpretable Failure Prediction for Vision-Language Models via Sparse Autoencoders

Jie Ma, Zongxi Liu, Yi Zhu

arXiv 2609.04276首次发表:更新:

发表机构

Wayne State University(韦恩州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于稀疏自编码器的FailSAE框架,将视觉语言模型故障预测建模为稀疏潜在激活分类任务,经三阶段故障感知训练提升性能与可解释性,优于基线方法且支持故障恢复。

AI 中文摘要

CLIP等视觉语言模型(VLMs)通过在共享嵌入空间中对齐视觉与文本表示,在多模态任务中取得了优异性能。随着VLMs在高风险领域的应用日益广泛,故障预测对于风险感知部署和人工干预至关重要。现有故障预测方法通常依赖置信度分数或辅助分类器,虽能有效预测VLMs故障,但可解释性有限。本研究探讨稀疏自编码器(SAEs)在VLMs可解释故障预测中的应用,将故障预测建模为稀疏SAE潜在激活上的分类任务,并引入三阶段故障感知训练流程,该流程在学习到的潜在方向保持可解释性的同时,提升对故障预测的信息性。实验结果表明,所提框架在故障预测性能上优于评估的基线方法。进一步分析显示,故障感知训练促使SAE潜在方向捕捉更多类特定概念;还利用SAE对模型表示在故障期间的变化进行概念级分析,揭示其从类特定概念向更模糊或风格相关概念转变;最后探究学习到的SAE潜在方向如何支持运行时故障恢复。

英文摘要

Vision-language models (VLMs), such as CLIP, have achieved strong performance across multimodal tasks by aligning visual and textual representations in a shared embedding space. As VLMs are increasingly used for high-stakes domains, failure prediction becomes critical for risk-aware deployment and human intervention. Existing failure prediction methods typically rely on confidence scores or auxiliary classifiers. Although these methods are effective on predicting VLM failures, they provide limited interpretability. In this work, we investigate the use of Sparse Autoencoders (SAEs) for interpretable failure prediction in VLMs. We formulate failure prediction as a classification task over sparse SAE latent activations and introduce a three-stage failure-aware training pipeline that encourages the learned latent directions to remain interpretable while becoming more informative for failure prediction. Our experiments show that the resulting framework outperforms the evaluated baselines in failure prediction. Further analysis suggests that failure-aware training encourages SAE latent directions to capture more class-specific concepts. We also use the SAE to provide a concept-level analysis of how model representations change during failures, revealing a shift from class-specific concepts toward more ambiguous or style-related concepts. Finally, we explore how the learned SAE latent directions can support runtime failure recovery.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑