arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24900cs.HC

与散文写作相比,代码写作期间大脑活动与大语言模型(LLM)嵌入之间的对齐性更强

Stronger Alignment between Brain Activity and LLM Embeddings during Code Writing compared to Prose Writing

Zachary Karas, Catie Chang, Kevin Leach, Yu Huang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过fMRI和VEMs发现,代码写作时大脑活动与LLM嵌入的对齐性显著强于散文写作,为相关AI系统设计提供了依据。

中文摘要 AI 辅助

编程是支撑现代软件系统的关键技能,但人们对代码写作背后的认知过程才刚刚开始了解,这限制了教育实践和开发者工具的发展。与此同时,大语言模型(LLM)越来越多地被用于辅助编程,而这些模型本身尚未被充分理解,可能会出现引入安全漏洞等不良行为。鉴于有证据表明LLM与大脑之间可能存在一些共享的认知表征,我们希望通过将这两个系统相互关联来增进对双方的理解。我们使用体素级编码模型(VEMs),将LLM嵌入与功能磁共振成像(fMRI)测量的自然写作任务期间的大脑活动关联起来。以23名参与者的按键输入作为提示,我们提取LLM嵌入来预测体素级血氧水平依赖(BOLD)信号,并通过预测信号与记录信号之间的相关性来量化对齐性。为了评估这种对齐性是否是编程特有的,还是可推广到其他生成过程,我们将代码写作与散文写作进行了比较。对齐性在右额极区域最强,且代码写作期间大脑活动由LLM嵌入预测的效果显著优于散文写作(p < 0.001,FDR校正)。在参与者内部,代码写作的最佳建模体素位置在LLM各层之间的一致性为66%,但在参与者之间的相似度差异很大(为39%)。我们的研究结果表明,在结构化代码生成期间,人类与LLM表征之间的对齐性更强,这对设计可预测代码生成但支持自然语言任务的AI系统具有启示意义。

英文摘要

Programming is a critical skill underlying modern software systems, yet the cognitive processes supporting code writing are only beginning to be understood, limiting educational practices and developer tools. At the same time, Large Language Models (LLMs) are increasingly used to assist programming. These models themselves are not well understood and can exhibit undesirable behavior like introducing security vulnerabilities. Given evidence that some cognitive representations may be shared between LLMs and the brain, we seek to improve our understanding on both fronts by relating these two systems to one another. We used Voxelwise Encoding Models (VEMs) to relate LLM embeddings to brain activity measured with functional Magnetic Resonance Imaging (fMRI) during naturalistic writing tasks. Using participants' (n = 23) keystrokes as prompts, we extracted LLM embeddings to predict voxelwise Blood Oxygen Level Dependent (BOLD) signal, quantifying alignment as the correlation between predicted and recorded signal. To assess whether this alignment is specific to programming or generalizes to other generative processes, we compared code writing to prose writing. Alignment was strongest in the right frontal pole, and brain activity was significantly better predicted by LLM embeddings during code writing than prose writing (p < 0.001, FDR-corrected). Within participants, the best-modeled voxel locations for code writing were 66% consistent across LLM layers but varied substantially between participants (39% similarity). Our findings suggest stronger alignment between human and LLM representations during structured code generation, with implications for designing AI systems that predict code generation but support natural language tasks.

↑