发表机构
AIM Intelligence; KT Corp.; Sionic AI; Lablup Inc.; Hankuk University of Foreign Studies(AIM智能公司; KT公司; Sionic人工智能公司; Lablup公司; 韩国外国语大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出难度校准的英韩谜题基准NOLLI,评估15种模型发现纯语言带来的性能差距小,文字系统密集型任务差距大,还揭示结构规模并非经验难度的可靠指标。
AI 中文摘要
我们推出NOLLI,这是一个程序生成的英韩谜题基准,旨在诊断韩国模型的性能差距产生的位置。它包含15种谜题类型(25个任务;7500个条目),每个实例都可通过种子重新生成,经确认具有唯一解,并采用确定性评分。我们没有将难度等同于规模更大,而是通过行为方式校准难度,对每个生成器进行调整,直到固定参考模型达到目标准确率区间。其三级设计结合了匹配的直接翻译、基于谚文音节子单元(jamo)的文字改编,以及基于韩国文化或正字法的纯韩语任务。我们评估了15种前沿、开放权重及韩国开发的模型;在整体准确率超过3%底线的12个模型中,匹配的英韩准确率在±10个百分点(TOST检验)范围内具有统计等价性,表明仅呈现语言带来的成本很小。文字系统密集型任务显示出更明显的差距:韩国密码(Korean Cipher)的准确率比英语低达68.7个百分点,而基于相同jamo的密码算术(Cryptarithmetic)则无系统性劣势,且jamo组合(Jamo Composition)的准确率可预测韩国密码的准确率。这些对比是诊断性而非因果性的,与多步骤子单元执行的难度一致。纯韩语任务将规则应用缺陷(符号各异)与亲属关系(Kinship)缺陷(在所有12个模型中均为正)区分开来。最后,15种类型中有7种的显著规模指标在简单到困难的任务中未增长,使结构规模成为经验难度的不可靠代理。
英文摘要
We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 puzzle types (25 tasks; 7,500 items), with every instance seed-regenerable, verified to have a unique solution, and scored deterministically. Rather than equating harder with bigger, we calibrate difficulty behaviorally, tuning each generator until a fixed reference model lands in target accuracy bands. Its three-level design combines matched direct translations, script adaptations over Hangul jamo (sub-syllabic letters), and Korean-only tasks grounded in Korean culture or orthography. We evaluate 15 frontier, open-weight, and Korean-developed models; among the 12 above a 3% overall-accuracy floor, matched English-Korean accuracy is statistically equivalent within a +/- 10 pp margin (TOST), suggesting little cost from presentation language alone. Writing-system-intensive tasks show sharper gaps: Korean Cipher falls behind English by up to 68.7 pp, whereas Cryptarithmetic over the same jamo shows no systematic penalty, and Jamo Composition accuracy predicts Korean Cipher accuracy. These contrasts are diagnostic rather than causal, consistent with difficulty in multi-step sub-syllabic execution. Korean-only tasks separate rule-application deficits, which vary in sign, from a Kinship deficit positive in all 12. Finally, a salient size measure fails to grow from Easy to Hard in 7 of 15 types, making structural size an unreliable proxy for empirical difficulty.