计算机认知科学的火花:来自模拟数据的理论可以泛化到人类
Sparks of In Silico Cognitive Science: Theories from Simulated Data Can Generalize to Humans
浏览论文内容
中文总结 AI 辅助
本研究通过AutoCog闭环系统在模拟器Centaur上发现行为理论,并验证其能泛化到人类数据,优于经典理论,表明不完美模拟器可拓宽理论搜索。
中文摘要 AI 辅助
行为基础模型已被提议作为跨情境的人类参与者替代品,但尚不清楚在它们上发现的理论是否能泛化到人类,还是仅仅表征了模拟器本身。我们运行了自动化认知科学家(AutoCog),这是一个闭环发现系统,其中LLM智能体设计区分理论的实验、收集响应、在竞争理论之间进行仲裁,并综合后继理论,整个过程完全基于由Centaur(一个人类行为基础模型)模拟的行为。在多属性决策情境中,AutoCog在Centaur上发现的理论泛化到了人类数据:它们在十个保留实验上优于经典理论,并且仅被通过在同一循环上对人类运行而发现的理论所匹敌。我们认为,尽管模拟器存在不可避免的缺陷,这一成功仍能实现,因为一个在竞争理论之间进行仲裁的发现循环对模拟器的要求低于估计。模拟器只需捕捉区分理论所需的规律性,而不必精确复现行为。因此,不完美的模拟器可以拓宽理论搜索范围,然后用人类数据来检验浮现的理论是否泛化。
英文摘要
Behavioral foundation models have been proposed as stand-ins for human participants across settings, but it is unclear whether theories discovered on them generalize to humans or merely characterize the simulator. We ran the Automated Cognitive Scientist (\textsc{AutoCog}), a closed-loop discovery system in which LLM agents design theory-discriminating experiments, collect responses, arbitrate between competing theories, and synthesize successors, entirely on behavior simulated by Centaur, a foundation model of human behavior. In a multi-attribute decision-making setting, the theories \textsc{AutoCog} found on Centaur generalized to human data: they outperformed canonical theories on ten held-out experiments and were rivaled only by theories found by running the same loop on people. We argue that this succeeds despite the simulator's inevitable imperfections because a discovery loop that arbitrates between competing theories demands less of its simulator than estimation does. The simulator only needs to capture the regularities that distinguish the theories, and not necessarily reproduce behavior precisely. Imperfect simulators can therefore widen the search over theories, with human data then testing whether the surfaced theories generalize.
发表机构
- Princeton University(普林斯顿大学)
- Cornell University(康奈尔大学)
- Helmholtz Munich(慕尼黑亥姆霍兹中心)
机构由 AI 辅助整理,请以论文原文为准。