arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更好的行为预测,更忠实的模型消融?来自序列选择的证据

Better Behavioral Prediction, More Faithful Model Ablations? Evidence from Sequential Choice

Hanbo Xie

arXiv 2609.36097首次发表:更新:

发表机构

School of Psychological and Brain Sciences; Georgia Institute of Technology(心理与脑科学学院; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过合成序列赌博机任务,检验了预测准确性是否保证消融响应的忠实性,发现准确预测器在捐赠者替换下响应可能远小于真实生成器,且模型排序因任务而异,强调需独立验证消融响应。

AI 中文摘要

使用预测模型来解释认知,需要的不仅仅是准确的行为预测。输入消融提供了一种有吸引力的途径:从模型中移除信息,并将由此产生的性能变化解释为该信息对行为重要性的证据。然而,这种推断假设模型对信息的依赖反映了产生行为的过程的依赖。我们在两个具有已知生成策略的合成序列赌博机任务中对此进行了测试,在这些任务中,当预测器无法获得反馈时,过去的选择可能仍然具有信息性。我们比较了从头训练的GRU和Transformer、微调的LLaMA模型以及认知模型,在系统变化的奖励贡献下。我们的分析区分了在没有奖励观察的情况下训练后的预测与固定预测器对捐赠者奖励替换的响应。出现了三个发现。首先,在动态任务中,没有奖励训练的神经模型对保留的选择的预测优于四个简单的训练拟合行为基线。其次,在匹配的捐赠者替换下,准确的预测器的响应可能远小于已知生成器。第三,在某些奖励权重下,神经网络的预测优于合并的强化学习模型,但选择概率的变化忠实度较低;模型排序在两个任务之间有所不同。这些独立测试结果将预测所需的信息与在序列选择中指定消融下的响应保真度区分开来。它们促使在将模型消融响应用于推断观察到的行为如何生成之前,独立于预测性能验证模型消融响应。

英文摘要

Using predictive models to explain cognition requires more than accurate behavioral predictions. Input ablations offer an appealing route: remove information from a model and interpret the resulting performance change as evidence of its importance for behavior. Yet this inference assumes that the model's dependence on information reflects the dependence of the process generating the behavior. We test it in two synthetic sequential bandit tasks with known generating policies, where past choices can remain informative when feedback is unavailable to a predictor. We compare GRUs and Transformers trained from scratch, a fine-tuned LLaMA model, and cognitive models across systematically varied reward contributions. Our analyses distinguish prediction after training without reward observations from the response of a fixed predictor to donor-reward replacement. Three findings emerge. First, in the restless task, neural models trained without rewards predict held-out choices better than four simple training-fitted behavioral baselines. Second, under matched donor replacement, accurate predictors can respond much less than the known generator. Third, at some reward weights, neural networks predict better than a pooled reinforcement-learning model but have less faithful changes in choice probabilities; the model ordering differs between the two tasks. These independent-test results separate information sufficient for prediction from response fidelity under a specified ablation in sequential choice. They motivate validating model-ablation responses independently of predictive performance before using them to infer how the observed behavior was generated.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑