小语言模型中的可解释跨语言对齐:探究日英双语大语言模型的文化与语用推理
Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs
浏览论文内容
中文总结 AI 辅助
本研究构建J-PragEval-v0基准,结合线性探针与教师强制评估探究TinySwallow-1.5B的日英语用表征,提出语用表征引导方法,下一步将扩展至Llama-3.1-Swallow-8B。
中文摘要 AI 辅助
大语言模型在英语上表现良好,但在与英语类型差异较大的语言上表现不佳,日语就是一个典型例子,目前对日语的评估仍依赖翻译质量和JGLUE风格基准,这类基准将词汇、句法和语用能力整合为单一分数。通用模型在日语用户的语用问题上表现不佳,包括敬语、内群体与外群体指代、语境敏感礼貌、零回指。本文提出J-PragEval-v0,这是一个最小对基准,从表面流畅性中分离出上述四种语用现象,并结合线性探针和教师强制对数概率评估,探究TinySwallow-1.5B(28层,隐藏层大小1536)中对应对比的位置。四种特征呈现三种分布:敬语语域清晰位于残差流中,第15层的平衡准确率达0.96,模型在93%的项目上会随情景翻转偏好的后续内容;隐性主语和内群体指代在最终提示词标记处无法线性解码(准确率分别为0.48和0.38),但翻转率分别为0.77和0.79,说明该对比是在生成过程中推导而非存储在提示词中;间接拒绝是负例,探针准确率达0.95,但在长度归一化教师强制下翻转率降至0.43,因为当前最小对将礼貌与后续内容长度混淆。本文还提出语用表征引导(Pragmatic Representation Steering),这是一种无参数推理时方法,可沿探针识别的类别均值差方向编辑残差流激活。其可行性通过间接论证而非直接展示:对比激活加法基线(即该方法将注入的相同几何结构)在存在线性信号的地方,能将探针准确率恢复至逻辑回归的1至2个点内。下一步将扩展至Llama-3.1-Swallow-8B。
英文摘要
Large language models work well on English and behave in poorly understood ways on languages typologically far from it. Japanese is a clean example, where evaluation still leans on translation quality and JGLUE-style benchmarks, which roll lexical, syntactic and pragmatic competence into a single score. The phenomena on which general-purpose models fail Japanese users are pragmatic: honorifics, in-group and out-group reference, context-sensitive politeness, zero anaphora. I introduce J-PragEval-v0, a minimal-pair benchmark isolating four such phenomena from surface fluency, and combine it with linear probes and teacher-forced log-probability evaluation to ask where inside TinySwallow-1.5B (28 layers, hidden size 1536) the corresponding contrasts live. The four features split three ways. Honorific register sits cleanly in the residual stream: 0.96 balanced accuracy at layer 15, and the model flips its preferred continuation with the scenario on 93 percent of items. Implicit subject and in-group reference are not linearly decodable at the final prompt token (0.48 and 0.38), yet flip rates are 0.77 and 0.79, so the contrast is worked out during generation rather than stored at the prompt. Indirect refusal is the negative case: 0.95 probe accuracy collapsing to a 0.43 flip rate under length-normalised teacher forcing, because the current minimal pairs conflate politeness with continuation length. I also specify Pragmatic Representation Steering, a parameter-free inference-time method that edits residual-stream activations along the class-mean-difference directions probing identifies. Feasibility is argued indirectly rather than demonstrated: the contrastive activation addition baseline, the same geometry the method would inject, recovers probe accuracy within one to two points of logistic regression wherever a linear signal exists. Scaling to Llama-3.1-Swallow-8B is the next step.
发表机构
- Sakaguchi–Inui Laboratory (Tohoku University)(东北大学坂口–乾实验室)
- Natural Language Understanding Team at RIKEN AIP(理化学研究所先进智能项目中心自然语言理解团队)
- Swallow Project at the Institute of Science Tokyo(东京科学大学Swallow项目)
- Sakana AI
- Tohoku University(东北大学)
- RIKEN AIP(理化学研究所先进智能项目中心)
- Institute of Science Tokyo(东京科学大学)
机构由 AI 辅助整理,请以论文原文为准。