发表机构
IPAI, Seoul National University; Dept. of ECE, Seoul National University(首尔国立大学人工智能研究所; 首尔国立大学电子与计算机工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多语言大语言模型中语言对指令层次结构合规性的影响,引入XIH-Bench基准,发现IH合规性有语言依赖不对称性及语言边界效应,还表明语言专业化带来多语言可靠性和安全风险。
AI 中文摘要
指令层次结构(IH)要求模型按来源对指令进行优先级排序,确保高优先级指令覆盖低优先级指令。尽管其对安全可控部署很重要,但现有评估几乎只关注英语,尚不清楚在多语言环境中IH合规性是否保持稳定。我们引入了XIH-Bench,这是一个用于多语言IH评估的基准,涵盖六种语言、四个领域和三种IH设置中的同语言和跨语言冲突。我们发现了两个一致的模式。首先,IH合规性表现出明显的语言依赖不对称性:在高优先级位置增强合规性的语言在低优先级位置可能具有破坏性。其次,跨语言冲突的合规性高于同语言冲突,我们将此现象称为语言边界效应。我们进一步表明,语言专业化会使模型偏好语言中的低优先级指令更难被覆盖,从而产生多语言可靠性和安全风险。
英文摘要
Instruction hierarchy (IH) requires models to prioritize instructions by source, ensuring that higher-priority instructions override lower-priority ones. Despite its importance for safe and controllable deployment, existing evaluations have focused almost exclusively on English, leaving it unclear whether IH compliance remains stable in multilingual settings. We introduce XIH-Bench, a benchmark for multilingual IH evaluation with both same-language and cross-language conflicts across six languages, four domains, and three IH settings. Across models, we find two consistent patterns. First, IH compliance exhibits a clear language-dependent asymmetry: a language that strengthens compliance in the higher-priority position can become disruptive in the lower-priority position. Second, cross-language conflicts yield higher compliance than same-language conflicts, a phenomenon we term the Language Boundary Effect. We further show that language specialization can make lower-priority instructions in model-favored languages harder to override, creating multilingual reliability and security risks.
CommentsAccepted to EMNLP 2026 (Main). Code and data are available at https://github.com/g1moon/Language-Shapes-IH