arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从大型语言模型中失语症图片命名错误轮廓中恢复损伤参数

Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models

Yong Yang, Roger Newman-Norlund, Xiang Guan, Saeed Ahmadi, Regan Willis, Nadra Salman, Kalil Warren, Sophie Arheix-Parras, Srihari Nelakuditi, Leonardo Bonilha, Christopher Rorden, Rutvik H. Desai, Julius Fridriksson

arXiv 2608.06429首次发表:更新:

AI 中文总结

该研究针对LLM的失语症图片命名错误轮廓,训练多任务神经网络恢复损伤参数,发现除层索引仅邻域可恢复外,修改百分比和噪声sigma可恢复,反事实验证达81.4%,还能泛化到中风幸存者数据,揭示Transformer层功能冗余。

AI 中文摘要

大型语言模型(LLMs)的可解释性方法描述其内部状态,但并未直接测试该状态是否在因果上足以产生观察到的行为。在早期工作中,我们对LLMs进行损伤处理,使其在图片命名任务中产生错误轮廓,而图片命名是评估失语症的核心任务,我们发现特定损伤会产生类似单个中风幸存者的错误。在此,我们提出逆问题:给定一个错误轮廓,能否恢复产生该错误的损伤参数,且该逆问题能揭示Transformer计算的什么特性?我们在4840种配置下,以层索引、修改百分比和噪声sigma对LLaVA-Vicuna 13B进行损伤处理,错误轮廓则通过七类临床分类法(正确、语义、无关、形式、混合、新语症、无反应)表征。我们训练了一个多任务神经网络,将错误轮廓映射回扰动参数。该问题可部分解决:在10个独立训练的逆模型中,修改百分比和噪声sigma可被恢复,而层索引仅能在邻域内被恢复。在反事实验证中,用恢复参数扰动的新模型实例在81.4%的案例中重现了目标行为。低层恢复与高反事实保真度之间的分离与Transformer层间的功能冗余一致,而标准可解释性方法未捕捉到该特性。作为分布外测试,我们将训练后的模型应用于278名中风幸存者的图片命名错误轮廓;恢复的参数具有综合征区分性,对扰动强度的区分性最强,表明其能泛化到训练分布之外。反事实验证为LLM可解释性主张提供了超越逆映射的通用框架。

英文摘要

Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient to produce the observed behavior. In earlier work, we lesioned LLMs to produce error profiles in picture naming, a central task for assessing aphasia, and found that specific lesions produced errors resembling those of individual stroke survivors. Here we ask the inverse question: given an error profile, can the lesion parameters that produced it be recovered, and what does this inverse problem reveal about transformer computation? Lesions in LLaVA-Vicuna 13B were parameterized by layer index, modification percentage, and noise sigma across 4,840 configurations, and error profiles were characterized by a seven-category clinical taxonomy (correct, semantic, unrelated, formal, mixed, neologism, no-response). We trained a multi-task neural network to map error profiles back to perturbation parameters. The problem admitted a partial solution: across 10 independently trained inverse models, modification percentage and noise sigma were recoverable, whereas layer index was recoverable only within a neighborhood. In counterfactual validation, a fresh model instance perturbed with the recovered parameters reproduced the target behavior in 81.4% of cases. This dissociation between low layer recovery and high counterfactual fidelity is consistent with functional redundancy across transformer layers, a property not captured by standard interpretability methods. As an out-of-distribution test, we applied the trained model to picture-naming error profiles from 278 stroke survivors; recovered parameters were syndrome-discriminative, most strongly for perturbation intensity, indicating generalization beyond the training distribution. Counterfactual validation provides a general framework for LLM interpretability claims beyond inverse mapping.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑