arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

揭示GAE的真实工作机制与固有缺陷:一种从经济学视角解释Transformer过程的方法

Unraveling the Real Working Mechanism and Inherent Flaws of GAE: A Method for Interpreting Transformer Processes from an Economic Perspective

Yongjin Cui, Xiaohui Fan

arXiv 2609.07213首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文从经济学视角揭示GAE解释方法的真实工作机制与固有缺陷,并提出融合过程与特征思想的CAH方法,用于更合理地解释Transformer模型。

AI 中文摘要

我们观察到一种现象:当前可解释人工智能领域的算法研究主要追求在若干代理指标上取得更好的性能。一方面,这些代理指标本身或多或少存在缺陷,无法恰当衡量方法的质量。另一方面,以指标为导向的研究方法往往导致对方法本身的合理性和可解释性的忽视。可解释人工智能缩写为XAI。以指标驱动的研究范式导致相关XAI方法本身缺乏可解释性。因此,需要对XAI方法进行可解释性研究,这可以戏称为XXAI。本文是我们关于XXAI的工作之一。本文以具有广泛影响力的模型解释方法——通用注意力模型可解释性(GAE)——为研究对象,该方法代表了重要的技术路线,探讨了该方法的真实工作机制和缺陷,以及它所代表的技术路线。基于本研究的结论,可能有必要重新审视或验证与GAE相关的方法及其领域应用。我们认为GAE是一种关注注意力过程的解释方法。在指出GAE的工作机制和缺陷后,我们提出了累积资产持有量(CAH),这是一种从经济零和博弈视角出发,融合基于过程和基于特征思想的更合理的Transformer解释方法。此外,值得注意的是,我们的方法适用于具有特殊标记的模型,而现有方法在此类模型中可能受到限制。本研究采用的模型简化研究方法和加法运算分析可能为XAI领域的其他研究工作提供启示。

英文摘要

We observe a phenomenon that current algorithmic research in the field of explainable artificial intelligence primarily pursues better performance on several proxy metrics. On the one hand, these proxy metrics themselves are more or less flawed and cannot properly measure the quality of methods. On the other hand, metric-oriented research approaches often lead to the neglect of the rationality and interpretability of the methods themselves. Explainable artificial intelligence is abbreviated as XAI. The metric-driven research paradigm has resulted in a lack of interpretability of the relevant XAI methods themselves. Accordingly, there is a need for interpretability research on XAI methods, which can be playfully referred to as XXAI. This paper is one of our works on XXAI. This paper takes Generic Attention-model Explainability (GAE), a widely influential model interpretation method , or rather, XAI method that represents an important technical route, as the research object, and explores the real working mechanism and flaws of this method as well as the technical route it represents. Based on the conclusions of this study, it may be necessary to re-examine or verify GAE-related methods and their domain applications. We argue that GAE is an interpretation method that focuses on the attention process. After pointing out the working mechanism and flaws of GAE, we propose Cumulative Asset Holdings (CAH), a more reasonable Transformer interpretation method integrating both process-based and feature-based ideas from an economic zero-sum games perspective. In addition, it is worth noting that our method is applicable to models with special tokens, where existing methods may suffer from limitations. The model simplification research method and the analysis of additive operations adopted in this study may provide inspiration for other research works in XAI.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑