MechSparse:机制引导的稀疏PEFT选择由任务形态决定
MechSparse: Mechanism-Guided Sparse PEFT Selection Is Task-Shaped
- RMIT University(皇家墨尔本理工大学)
- Ho Chi Minh City University of Technology (HCMUT), VNU-HCM(胡志明市理工大学(HCMUT),越南国立大学胡志明市分校)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究提出MechSparse机制引导的稀疏PEFT选择方法,通过激活修补评分定位关键头与MLP块,在信息抽取和机器翻译任务上验证其效果,发现激活范数在结构主导任务中更优,因果选择器适合内容主导任务。
中文摘要 AI 辅助
机制可解释性识别出携带特定行为的注意力头和MLP块的稀疏子集。我们探究这些因果信号是否比从业者已使用的廉价启发式方法更有效地指导小规模PEFT预算的放置位置。\u003cmethod\u003e通过干净/损坏探针上的归一化激活修补恢复度对注意力头和MLP块进行评分,并仅在选定位置训练LoRA/QLoRA;\u003cmethodc\u003e为小的层内联合子集添加有界信用。我们在Ministral-8B/NF4上,于三个实验单元中比较随机、幅度、激活范数及梯度/Fisher方法:斯瓦希里语跨度-JSON信息抽取(IE)在b=0.25%和1.0%预算下,以及英语到斯瓦希里语机器翻译(MT)在b=1.0%预算下。因果选择器从未在主要指标上获胜。在头条IE实验单元(3个种子,600个预测上的配对自助置信区间)中,\u003cmethodc\u003e比随机方法高出+0.079的跨度+类型F1,比梯度/Fisher高出+0.174,但落后激活范数0.028,且具有最小的跨种子标准差(±0.003)。在MT上,所有四个选择器均在0.30 BLEU范围内,且每个配对置信区间包含零。模式与跨度分解解释了IE差距:激活范数捕获刚性JSON例程,而因果分数跟踪内容敏感位置。我们提炼出初步诊断——当输出结构占主导时偏好激活范数,将因果选择器视为内容主导任务的假设——并发布掩码、分数、预测和评估文件以供直接复现。
英文摘要
Mechanistic interpretability identifies sparse subsets of heads and MLP blocks that carry specific behaviors. We ask whether such causal signals can guide where to place a small PEFT budget more effectively than the cheap heuristics practitioners already use. \method{} scores attention heads and MLP blocks by normalized activation-patching recovery on clean/corrupted probes and trains LoRA/QLoRA only on the selected sites; \methodc{} adds bounded credit for small within-layer joint subsets. We compare against random, magnitude, activation-norm, and gradient/Fisher on Ministral-8B/NF4 in three cells: Swahili span-JSON information extraction (IE) at $b{=}0.25\%$ and $1.0\%$, and English$\to$Swahili machine translation (MT) at $b{=}1.0\%$. The causal selectors never win the primary metric. On the headline IE cell (3 seeds, paired-bootstrap CIs over $600$ predictions), \methodc{} beats random by $+0.079$ span+type F1 and gradient/Fisher by $+0.174$, but trails activation-norm by $0.028$, with the smallest cross-seed std ($\pm 0.003$). On MT all four selectors lie within $0.30$ BLEU and every paired CI includes zero. A schema-versus-span decomposition explains the IE gap: activation-norm captures the rigid JSON routine, while causal scores track content-sensitive sites. We distill a preliminary diagnostic -- prefer activation-norm when output structure dominates, treat causal selectors as a hypothesis for content-dominated tasks -- and release masks, scores, predictions, and evaluation files for direct replay.