arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DMDIntel:通过动态模态分解解释大型语言模型

DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition

Amogh Joshi, Animesh Mukherjee, Sergey Utyuzhnikov

arXiv 2608.13048首次发表:更新:

发表机构

IIT Kharagpur; University of Manchester(印度理工学院克勒格布尔分校; 曼彻斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出DMDIntel方法,通过动态模态分解为大型语言模型分类任务构建输入归因流程,经多数据集与模型系列实验验证,其归因性能优于主成分分析、集成梯度、SHAP等现有最优技术。

AI 中文摘要

本研究提出DMDIntel,其利用动态模态分解(DMD)使大型语言模型(LLM)在分类任务中的预测可解释。该方法构建了输入归因流程:首先将LLM的隐状态分解为显著模式(即模态),再基于输入词元在这些模态上的投影值为其分配排名。在三个数据集和三个模型系列上开展的严谨实验一致表明,采用DMDIntel得到的输入词元排名归因结果,显著优于主成分分析、集成梯度和SHAP等现有最优技术。

英文摘要

In this work, we introduce DMDIntel which uses dynamic mode decomposition (DMD) to make the predictions made by LLMs in a classification task interpretable. It develops an input attribution pipeline, that first decomposes the hidden states of an LLM into prominent patterns, also known as modes, and then associates ranks to the input tokens based on the projection values on those modes. Rigorous experiments across three datasets and three model families consistently show that the ranked attribution of input tokens obtained using DMDIntel by far outperforms state-of-the-art techniques such as principal component analysis, integrated gradients and SHAP.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑