人类认知与行为的小型基础模型
Small Foundation Models of Human Cognition and Behaviour
浏览论文内容
中文总结 AI 辅助
该研究训练不同规模的小型认知微调模型,发现其在分布内表现稳定、分布外泛化需更大模型,且可作为心理学实验噪声上限估计器。
中文摘要 AI 辅助
基于人类行为数据微调的大型语言模型已成为通用认知代理,但所需的规模以及这些模型是处理任务结构还是利用统计捷径仍是未解决的问题。我们在Psych-101数据集上训练了14个模型,参数规模从1.35亿到140亿,涵盖四个架构系列;Psych-101是包含160个实验的1070万次试次级选择的数据集。在分布内场景中,规模几乎不产生影响,模型落在狭窄区间内,仿佛存在一个上限,0.6亿至10亿参数足以匹配700亿参数基线在保留参与者上的表现。在分布外场景中,该区间形成明显更陡峭的缩放梯度,更大的模型在泛化到新任务结构时具有明显优势。为确定这些模型使用的信息,我们进行两项诊断:在27个实验中逐步移除四个提示通道——任务指令、实验刺激、结果反馈和选择历史,并打乱试次顺序。屏蔽刺激和反馈内容会破坏75.7%的已学习信息,使模型表现低于随机水平,表明仅靠选择历史无法解释性能。打乱顺序显示,在试次独立的任务上模型表现出不变性,但在试次顺序由先前反应决定的任务上具有敏感性。因此,经过认知微调的小型模型有望作为心理学实验的噪声上限估计器,不过其范围仍受限于训练中所见的范式。
英文摘要
Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. For in-distribution simulations, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.