arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MURANO:将机械可解释性实验设计、运行与复现作为可组合流水线

MURANO: Design, Run, and Reproduce Mechanistic Interpretability Experiments as Composable Pipelines

Alireza Bayat Makou, Emirhan Böge, Phu Gia Hoang, Federico Tiblias, Jingcheng Niu, Subhabrata Dutta, Richard Eckart de Castilho, Iryna Gurevych

arXiv 2608.30662首次发表:更新:

发表机构

Technical University of Darmstadt; Cluster of Excellence “Reasonable Artificial Intelligence” (RAI); National Research Center for Applied Cybersecurity ATHENE; Zuse School ELIZA; Ubiquitous Knowledge Processing (UKP) Lab(达姆施塔特工业大学; “合理人工智能”卓越集群; 国家应用网络安全研究中心ATHENE; 楚思学校ELIZA; 普适知识处理实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出开源框架Murano,将大语言模型机械可解释性研究的五类操作封装为可组合步骤,可复现相关研究并开展案例分析,解决了多库适配的问题。

AI 中文摘要

本文面向各学科研究者,提出了开源框架Murano,用于设计、运行和复现大语言模型的机械可解释性研究。这类研究通常结合加载、记录、归因、干预与评估操作,而现有库往往仅关注该工作流的不同部分,导致使用多个库的研究者需要适配某一库的输出以用于另一库。为弥合这一差距,Murano将上述五个领域的操作表示为可组合步骤,各步骤交换命名结果构件,并声明所需输入与产生的输出;流水线按指定顺序执行步骤,且当组件标识在操作间传递时,Murano使用规范地址。Murano基于现有可解释性与机器学习库构建,本文通过复现两项已发表的可解释性研究,以及一项说明性的稀疏自编码器案例研究,对Murano进行了演示。

英文摘要

This paper presents Murano, an open source framework for designing, running, and reproducing mechanistic interpretability studies of large language models, intended for researchers across disciplines. These studies often combine loading, recording, attribution, intervention, and evaluation, while existing libraries tend to focus on different parts of this workflow. As a result, researchers using several libraries may need to adapt outputs from one for use by another. To bridge this gap, Murano represents operations from these five areas as composable steps. Steps exchange named result artifacts and declare the inputs they require and the outputs they produce. A pipeline executes its steps in the order supplied, and Murano uses canonical addresses when component identities pass between operations. Murano builds on existing interpretability and machine learning libraries. We demonstrate Murano through two reproductions of established interpretability studies and one illustrative sparse autoencoder case study.

CommentsAccepted to the EMNLP 2026 System Demonstrations Track. 11 pages, 6 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑