arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

小语言模型在抽象推理任务上的系统性研究

A Systematic Study of Small Language Models on Abstract Reasoning Tasks

Nur A Zarin Nishat, Jens Lehmann, Andrei Aioanei, Sahar Vahdati

arXiv 2610.08680首次发表:更新:

发表机构

Leibniz University of Hanover; TIB – Leibniz Information Centre for Science and Technology and University Library; Amazon(莱布尼茨汉诺威大学; TIB – 莱布尼茨科学技术信息中心与大学图书馆; 亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究系统剖析小语言模型在抽象推理任务上的表现,发现分布内准确率可达但习得敏感且分布外性能骤降,模型、适应机制、评估分布和响应格式均影响推理分数。

AI 中文摘要

在抽象推理基准上的端点准确率并不能揭示语言模型是习得了可迁移的规则还是拟合了特定于分布的规律。我们在ARC-TGI基准上研究小语言模型中的这一区别,该基准将抽象网格变换组织成可控的任务族,并支持重采样、空间平移和跨基准迁移。在超过1000次运行中,我们在监督微调下对仅解码器、编码器-解码器和混合专家模型族进行了剖析。我们考察了技能习得的效率和稳定性、训练分布之外的鲁棒性、与模型族和任务表述的交互,以及伴随行为差异的逐层注意力特征。可实现可观的分布内准确率,但习得过程对优化敏感且在任务族间分布不均。在训练分布之外性能急剧下降,包括规则保留但网格尺度变化的情况。更大的训练集深度和广度带来不均衡的收益,而额外上下文示例的效果取决于模型族。可执行规则归纳也产生了直接网格生成下未观察到的正确解决方案。在选定任务上,注意力诊断显示出不同的集中度和上下文依赖特征,但未建立普遍的因果机制。总体而言,抽象推理分数取决于模型、适应机制、评估分布和响应格式。

英文摘要

Endpoint accuracy on abstract-reasoning benchmarks does not reveal whether a language model has acquired a transferable rule or fit distribution-specific regularities. We study this distinction in small language models on the ARC-TGI benchmark, which organizes abstract grid transformations into controllable task families and supports resampling, spatial shifts, and cross-benchmark transfer. Across more than 1,000 runs, we profile decoder-only, encoder--decoder, and mixture-of-experts model families under supervised fine-tuning. We examine the efficiency and stability of skill acquisition, robustness beyond the training distribution, interactions with model family and task formulation, and layer-wise attention signatures that accompany behavioral differences. Substantial in-distribution accuracy is attainable, but acquisition is sensitive to optimization and unevenly distributed across task families. Performance deteriorates sharply outside the training distribution, including when the rule is retained but grid scale changes. Greater training-set depth and breadth yield uneven gains, while the effect of additional in-context examples depends on model family. Executable-rule induction also yields correct solutions not observed under direct grid generation. On selected tasks, attention diagnostics show distinct concentration and context-dependence profiles, but do not establish general causal mechanisms. Overall, abstract-reasoning scores are conditional on the model, adaptation regime, evaluation distribution, and response format.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑