用人类认知基础模型预测程序理解能力
Predicting Program Comprehension with Foundation Models of Human Cognition
浏览论文内容
中文总结 AI 辅助
研究探讨用人类认知基础模型预测程序理解能力,通过评估基于160个心理实验训练的Centaur在9个程序理解研究中的表现,发现其比基础模型更契合人类响应模式,为软件工程中开发者行为建模提供了新思路。
中文摘要 AI 辅助
软件工程依赖开发者理解代码的能力,尽管经过数十年研究,但预测他们如何理解代码仍是一个开放挑战。现有方法要么依赖限制准确性的简化代理措施,要么依赖需要复杂实验设置且难以扩展和实际应用的非平凡测量。心理学近期工作提出了另一种观点:可通过从大规模行为数据中学到的认知规律来捕捉人类行为。本文在程序理解背景下探索此观点。我们在9个先前发表的程序理解研究中评估了基于160个一般心理实验训练的基础模型Centaur,评估其预测响应分布与人类响应数据的匹配程度,并将Centaur的性能与其基础模型Llama 3.1进行比较。为更好理解其性能来源,我们进行了消融研究。简而言之,我们发现Centaur比其基础模型更接近人类响应模式,显著更少依赖先前试验和响应的信息,且从任务相关信息中受益更多。这些发现表明从一般心理数据中学到的行为模式可转移到诸如程序理解等复杂软件工程任务中。更广泛地说,它们指出人类认知基础模型可作为软件工程中开发者行为建模的基础。
英文摘要
Software engineering depends on the ability of developers to understand code, yet predicting how they do so remains an open challenge despite decades of research. Existing approaches rely either on simplified proxy measures that limit accuracy or on non-trivial measurements requiring elaborate experimental setups that are difficult to scale and apply in practice. In contrast, recent work in psychology suggests an alternative perspective: Instead of modeling task-specific phenomena directly, human behavior can be captured through cognitive regularities learned from large-scale behavioral data. This idea treats complex human behavior as the observable outcome of underlying cognitive processes that manifest consistently across tasks and domains. In this paper, we explore this perspective in the context of program comprehension. We evaluate Centaur, a foundation model trained on 160 general psychological experiments, on 9 previously published program-comprehension studies. We assess how well its predicted response distributions align with human response data and compare Centaur's performance to its base model, Llama 3.1. To better understand the source of its performance, we conduct ablation studies to isolate the contribution of different sources of information, such as the code artifacts, task-related context, and prior trials and participant responses. In a nutshell, we find that Centaur more closely aligns with human response patterns than its base model, is significantly less reliant on information from prior trials and responses, and benefits more from task-related information. These findings suggest that behavioral patterns learned from general psychological data can transfer to complex software engineering tasks such as program comprehension. More broadly, they point toward foundation models of human cognition as a basis for modeling developer behavior in software engineering.