ADAGE:一种用于类比推理评估的语言无关管道
ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation
浏览论文内容
中文总结 AI 辅助
研究针对多语言推理评估依赖翻译基准的问题,提出ADAGE语言无关管道,结合母语者策划与大语言模型辅助生成构建基准,通过为多种语言构建基准并评估模型,发现文化推理差距,还发布了相关内容。
中文摘要 AI 辅助
多语言推理评估大多依赖翻译英语基准,这会引入语言假象且无法测试基于文化的推理。我们引入了ADAGE(基于设计的类比难度评估以进行有根据的评估),这是一种语言无关管道,它将母语者策划与大语言模型辅助生成相结合,为抽象类比推理构建具有挑战性、无需翻译的基准。我们通过为阿拉伯语、阿姆哈拉语和日语构建基准来验证ADAGE。评估14个开放权重模型时,我们发现了一致的文化推理差距:在英语谚语推理中表现良好的模型在所有三个母语基准上都大幅挣扎,准确率相对于英语下降了12 - 52个百分点。我们发布了该管道、所有三个基准和完整评估套件。
英文摘要
Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts and fails to test culturally-grounded reasoning. We introduce ADAGE (Analogical Difficulty-by-design Assessment for Grounded Evaluation), a language-agnostic pipeline that combines native-speaker curation with LLM-assisted generation to construct challenging, translation-free benchmarks for abstract analogical reasoning. We validate ADAGE by constructing benchmarks for Arabic, Amharic, and Japanese. Evaluating 14 open-weight models, we find a consistent cultural reasoning gap: models that perform well on English proverb reasoning struggle substantially on all three native benchmarks, with accuracy dropping by 12--52 percentage points relative to English. We release the pipeline, all three benchmarks, and the full evaluation suite.
发表机构
- Haverford College(哈弗福德学院)
机构由 AI 辅助整理,请以论文原文为准。