arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2607.03936cs.CL

方言能否像语言一样被引导?阿拉伯语大语言模型中的稀疏神经元和分布式方向

Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs

  • Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
  • Qatar Computing Research Institute, Hamad Bin Khalifa University(哈马德·本·哈利法大学卡塔尔计算研究所)
  • Tohoku University(东北大学)
  • RIKEN(理化学研究所)

机构由 AI 辅助整理,请以论文原文为准。

Kareem Elozeiri, Mervat Abassy, Omar Kallas, Fahim Dalvi, Preslav Nakov, Kentaro Inui, Nadir Durrani

AI总结:

研究阿拉伯语NLP中方言数据稀缺致模型生成问题,通过神经元层面分析和向量引导方法,揭示方言知识几何结构,提供无需微调的方言控制框架。

AI中文摘要:

阿拉伯语NLP的一个关键挑战是方言数据相对于现代标准阿拉伯语(MSA)的稀缺,这导致大语言模型过度生成MSA并在方言准确生成方面存在困难。从可解释性角度来看,这提出了一个基本问题:方言特征在模型内部何处以及如何编码,以及这些表示能否在不进行微调的情况下用于改进方言生成?本研究调查了两种互补的推理时方法,它们同时作为可解释性探针和控制机制。首先,我们进行了神经元层面的分析,识别出编码特定方言特征的稀疏神经元群体,并表明放大或抑制这些神经元可以将模型输出导向目标方言。其次,受单神经元层面方言特征纠缠的启发,我们应用了一种向量引导方法,该方法提取特定方言的激活方向并在推理期间注入它们。总之,这些方法阐明了阿拉伯语大语言模型中方言知识的几何结构,并提供了一个有原则的、基于可解释性的方言控制框架,而无需特定于方言的微调。

英文摘要:

Dialectal data are scarce relative to Modern Standard Arabic (MSA), causing Arabic LLMs to overproduce MSA and struggle with dialectally accurate generation. This raises a fundamental interpretability question about where and how dialectal features are encoded within model internals and whether these representations can improve dialect generation without fine-tuning. We study two inference-time approaches as interpretability probes and control mechanisms. First, neuron-level analysis identifies sparse populations that encode dialect-specific features and tests whether amplifying or suppressing them steers model outputs toward target dialects. Second, vector steering extracts dialect-specific activation directions and injects them during inference, motivated by feature entanglement at the neuron level. We find that these neurons are real but only partially explanatory. They occupy under 1\% of MLP dimensions but span only 5\% to 21\% of the residual dialect direction. This limited coverage is causally consequential. Neuron steering reinforces dialect in some varieties when the prompt is already dialectal but cannot induce it from MSA prompts, whereas vector steering succeeds in both settings. Arabic dialects are therefore steerable mainly through distributed rather than localized representations

↑