arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

骨架应该说什么语言?多语言推理中的语言选择

Which Language Should a Skeleton Speak? Language Choices in Multilingual Reasoning

HyeonSeok Lim, SeungWoo Song, Inho Won, Hoyun Song, Jihyo Kim, KyungTae Lim

arXiv 2610.09607首次发表:更新:

发表机构

ETRI; KAIST; Dankook University(韩国电子通信研究院; 韩国科学技术院; 檀国大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出LASEF框架,探索多语言数学推理中骨架语言的选择,发现英语骨架平均有微小正向效果但非普遍最优,效应呈三种模式,需多层面探索。

AI 中文摘要

基于骨架的推理提示是一种有前景的免训练方法,用于结构化大语言模型的推理,但先前的工作主要假设以英语为中心的环境。我们提出了语言感知骨架探索框架(LASEF),以研究多语言数学推理中的骨架语言选择。跨数学基准、模型规模及语言,我们表明英语骨架平均产生较小的正向趋势,在较小模型和低资源语言中最为明显。然而,在修正后,很少有语言层面的增益保持显著,且英语并非普遍最优。结合贪婪解码、多轮评估、翻译消融和跨基准验证,我们进一步发现骨架语言效应的三种模式:方向一致、评估与基准依赖,以及不对称负效应。这些效应不能仅由生成质量完全解释。总体而言,骨架语言是一个依赖上下文的设计变量,需要多层面的探索。所有资源已在此https URL发布。

英文摘要

Skeleton-based reasoning prompting is a promising training-free approach for structuring LLM reasoning, but prior work largely assumes an English-centric setting. We propose the Language-Aware Skeleton Exploration Framework (LASEF) to study skeleton-language choice in multilingual mathematical reasoning. Across math benchmarks, model scales, and languages, we show that English skeletons yield a small positive tendency on average, most visible for smaller models and low-resource languages. However, few language-level gains remain significant after correction, and English is not universally optimal. Combining greedy decoding, multi-rollout evaluation, translation ablation, and cross-benchmark validation, we further find three patterns of skeleton-language effects: directionally consistent, evaluation- and benchmark-dependent, and asymmetric negative. These effects cannot be fully explained by generation quality alone. Overall, skeleton language is a context-dependent design variable that requires multi-level exploration. All resources are released at https://github.com/lhsstn/LASEF.

CommentsAccepted to EMNLP 2026 (Findings)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑