arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13072cs.LGcs.AIcs.CL

MAxBench:一种多项式概念恢复基准

MAxBench: A Multinomial Concept Recovery Benchmark

Divya Appapogu, Freya Behrens, Yonatan Belinkov, Aaron Mueller

首次发表
浏览论文内容

中文总结 AI 辅助

MAxBench是一个与几何无关的多项式概念恢复基准,通过比较10种定位方法,发现仿射子空间引导更可靠且召回率更高,且优势源于非零偏移,流形引导有竞争力,但无方法优于提示。

中文摘要 AI 辅助

对语言模型行为的细粒度控制(例如,引导)是可解释性研究中更具可操作性的成果之一。对于诸如拒绝之类的二元概念,激活空间中的单个方向通常足以实现引导。然而,许多概念并非二元的:动物和国家包含许多子类别,每个子类别又包含多个实例。对于这些概念,可能的表示几何形状的搜索空间远大于二元概念;因此,尚不清楚哪些几何形状最合适,也不清楚哪些方法最有效地恢复它们。在这项工作中,我们引入了MAxBench,一个基于从恢复的概念表示中采样的、与几何无关的多项式概念表示评估框架。我们使用MAxBench比较了6个概念和4个模型上的10种定位方法(涵盖5种几何类型)。利用该框架,我们发现:(i)仿射子空间比秩一或线性子空间引导更可靠且召回率更高;(ii)这一优势很大程度上归因于更好的非零偏移,而非基的选择;(iii)流形引导在适用时与最佳方法具有竞争力;(iv)没有方法始终优于提示(prompting),这与先前关于二元概念的发现一致。这些发现强调了将可解释性研究和元评估的范围扩展到具有更多样化结构的概念的重要性。

英文摘要

Fine-grained control of language model behaviors (e.g., steering) is among the more actionable outcomes of interpretability research. For binary concepts such as refusal, a single direction in activation space often suffices for steering. However, many concepts are not binary: Animals and Countries contain many subcategories, each with multiple instances. For these concepts, the search space over possible representation geometries is far larger than for binary concepts; it is thus not clear what geometries are most appropriate, nor what methods are most effective at recovering them. In this work, we introduce MAxBench, a geometry-agnostic evaluation framework for multinomial concept representations based on sampling from the recovered concept representation. We use MAxBench to compare 10 localization methods (covering 5 geometry types) across 6 concepts and 4 models. Using this framework, we find that (i) affine subspaces steer more reliably and have greater recall than rank-one or linear subspaces; (ii) much of this advantage is due to better non-zero offsets rather than the choice of bases; (iii) manifold steering is competitive with the best methods when applicable; and (iv) no method consistently outperforms prompting, in alignment with prior findings on binary concepts. These findings underscore the importance of expanding the scope of interpretability research and meta-evaluation to concepts with more varied structure.

发表机构

  • Boston University(波士顿大学)
  • Technion – Israel Institute of Technology(以色列理工学院)
  • Kempner Institute, Harvard University(哈佛大学肯普纳研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑