基于减法映射的扰动型区域可解释性方法(PRISM):语言模型与卒中后失语症中的命名错误分离现象
Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia
浏览论文内容
中文总结 AI 辅助
本研究开发PRISM方法,将神经成像减法分析适配至Transformer语言模型,对比模型与213名卒中后失语症患者的命名错误模式,验证了音素偏向性分离等结果,为模型功能专业化提供可证伪的空间分辨率测试工具。
中文摘要 AI 辅助
大型语言模型的机制可解释性缺乏空间分辨率高、可证伪的工具,用于测试内部组件是否专门负责不同的认知操作。我们将人类神经成像的标准框架——减法分析从生物大脑适配到受扰动的Transformer模型,并对两种底物(生物大脑与受扰动的Transformer)并行应用相同逻辑。基于脑-大语言模型统一模型(BLUM)(该模型显示,经层扰动的LLaVA-1.6-Vicuna-13B的错误谱与失语症患者的损伤模式匹配),我们开发了PRISM(基于减法映射的扰动型区域可解释性方法)。PRISM映射费城命名测试的7个临床类别,成对减去错误类别,并将每个扰动种子视为组分析中的一个受试者,沿层轴进行无阈值集群增强。我们对213名慢性卒中后失语症患者进行结构匹配分析,采用相关差异损伤-症状映射,并在保留的拆分数据上重复了两侧结果。两种设计在受试者维度(种子、患者)、空间维度(层、脑图谱分区皮层)和阈值设置上匹配,但对比算子不同:大语言模型采用受试者内错误比例差异,皮层采用受试者间相关差异。两种底物均恢复了稳健的音素偏向性分离、一个深层集群和一个额-岛盖皮层集群,且均得到重复验证;语义偏向性方向在两侧均为符号一致但无统计学意义的趋势。因此,PRISM为Transformer语言模型中功能专业化主张提供了可证伪、空间分辨率高的测试。确认性的感兴趣区域(ROI)水平干预(PRISM第3阶段)(其支持最强的因果机制主张)留待后续工作开展。
英文摘要
Mechanistic interpretability of large language models lacks spatially resolved, falsifiable tools for testing whether internal components are specialized for distinct cognitive operations. We adapt subtraction analysis, the standard framework of human neuroimaging, from biological brains to perturbed transformers, and apply the same logic to both substrates in parallel. Building on the Brain-LLM Unified Model (BLUM), which showed that layer-perturbed LLaVA-1.6-Vicuna-13B error profiles match the lesion patterns of aphasic patients, we develop PRISM (Perturbation-based Regional Interpretability through Subtraction Mapping). PRISM maps the seven clinical Philadelphia Naming Test categories, subtracts error classes pairwise, and treats each perturbation seed as a subject in a group analysis with threshold-free cluster enhancement along the layer axis. We run a structurally matched analysis on 213 chronic post-stroke aphasia patients using correlation-difference lesion-symptom mapping, and replicate both sides on held-out splits. The designs match in subject dimension (seeds, patients), spatial dimension (layers, atlas-parcellated cortex) and thresholding, but the contrast operator differs: a within-subject error-proportion difference for the LLM, a between-subject correlation difference for the cortex. Both substrates recover a robust phonemic-favoring dissociation, a deep layer cluster and a frontal-perisylvian cortical cluster, both replicating; the semantic-favoring direction is a consistently signed but non-significant trend on both. PRISM thus gives a falsifiable, spatially resolved test of functional-specialization claims in transformer language models. A confirmatory ROI-level intervention (PRISM Stage 3) licensing the strongest causal-mechanism claim is left to subsequent work.
发表机构
- University of South Carolina(南卡罗来纳大学)
- ALLT.AI, LLC(ALLT.AI有限责任公司)
- USC School of Medicine(南卡罗来纳大学医学院)
机构由 AI 辅助整理,请以论文原文为准。