发表机构
University of Lagos(拉各斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探究非洲语言NLI自适应推断的基于表示的难度估计,发现内部表示统计量无法为该场景提供有用的示例级难度信号,且自适应路由的表现不如始终高开销的推断。
AI 中文摘要
我们探究内部表示统计量是否能为多语言非洲自然语言处理(NLP)的自适应推断提供有用的示例级难度信号,结果发现在该场景下无法实现这一点。我们使用冻结的现成检查点研究了15种非洲语言的自然语言推理,报告了四项结果:第一,AfriXNLI的英语配置在1050个示例中有1047个与XNLI评估数据完全一致,某广泛使用的NLI检查点在该测试拆分上的得分为1.000,这与该检查点接触过XNLI测试数据的情况一致;由于AfriXNLI源自XNLI,其英语、法语和斯瓦希语配置无法作为XNLI训练模型的干净评估数据。第二,参数数量无法可靠地对非洲语言的能力进行排序:我们的更大检查点在7种语言中表现更好,在8种语言中表现更差,无显著总体差异。第三,在三种多语言表示空间中,角离散度始终比有效秩更具语言决定性,因此池化相关性可能会放大其中一个指标并掩盖另一个指标。第四,经语言控制后保留的关联取决于目标:有效秩可预测升级带来的概率增益,但无法预测升级是否会改变预测结果,而廉价模型的置信度则呈现相反模式;两个目标的相关性仅为0.655。在所测试的模型、信号和计算预算下,没有任何评估信号能使自适应路由优于始终高开销的推断,尽管神谕在60%的计算量下比后者高出11个准确率点。我们的核心方法论发现是:一个表示统计量对于某一计算收益概念可能具有统计显著性,但对于另一概念却无关,因此不适合作为决策变量。
英文摘要
We ask whether internal representation statistics can provide useful example-level difficulty signals for adaptive inference in multilingual African NLP, and find that they cannot in this setting. Studying natural language inference across 15 African languages with frozen off-the-shelf checkpoints, we report four results. First, AfriXNLI's English configuration shares 1,047 of its 1,050 examples verbatim with XNLI evaluation data, and one widely used NLI checkpoint scores 1.000 on that test split, consistent with XNLI test exposure. Because AfriXNLI is derived from XNLI, its English, French and Swahili configurations cannot serve as clean evaluations for XNLI-trained models. Second, parameter count does not reliably order capability across African languages: our larger checkpoint is better in seven languages and worse in eight, with no significant aggregate difference. Third, across three multilingual representation spaces, angular dispersion is consistently more language-determined than effective rank, so pooled correlations can inflate one and mask the other. Fourth, the association that survives language control depends on the target: effective rank predicts probability gain from escalation but not whether escalation changes the prediction, while cheap-model confidence shows the opposite pattern; the two targets correlate at only 0.655. Under the tested models, signals, and compute budgets, no evaluated signal makes adaptive routing preferable to always-expensive inference, although an oracle exceeds it by 11 accuracy points at 60% of the compute. Our central methodological finding is that a representation statistic can be statistically significant for one notion of computational benefit while being irrelevant to another, and therefore be a poor decision variable.
Comments21 pages, 3 figures, 10 tables. Submitted to MIRG-ICAIR 2026