发表机构
University of Michigan; Sandia National Laboratories; University of Tennessee, Knoxville(密歇根大学; 桑迪亚国家实验室; 田纳西大学诺克斯维尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过调查527名科研人员,发现AI编程助手主要用于数据处理等五类任务,验证方式以个人运行为主,且信心与经验相关,建议界面支持任务适配的评估。
AI 中文摘要
生成式人工智能已进入研究编程领域,但关于研究人员将哪些任务交给它,以及他们如何判断其代码是否正确,目前证据甚少。我们基于2025年一项针对编写代码的研究人员(其中大多数在美国大学工作)的调查,收集了527份自由文本回答。在每份回答中,一位研究人员叙述了其工作中的单一任务、使用AI工具的方式以及评估结果所采取的行动。我们对所报告的任务和评估策略进行了编码,并将两者与编程经验、研究领域和信心评级相关联。使用集中在五项任务:数据处理、可视化、调试、数学/科学计算和统计分析。评估是非正式且个体化的。超过一半的描述提到运行生成的代码,而自动化测试和他人审查则很少见。使用场景和评估策略随编程经验变化不大,但信心却不同:经验较少的程序员更信任AI而非自己,而经验丰富的程序员则相反。评估信心与所报告的策略无关,其最强的相关因素是对工具和自身的信心。验证AI对科学代码的贡献主要依赖于个人判断,而非共享的测试或审查基础设施。界面应支持适合任务的评估,而不是将其留给用户。
英文摘要
Generative AI has entered research programming, yet there is little evidence about which tasks researchers hand to it or how they decide whether its code is correct. We draw on 527 free-text responses to a 2025 survey of researchers who write code, most of them at U.S. universities. In each response, a researcher recounts a single task from their own work, the way they used an AI tool for it, and what they did to assess the result. We coded the task and the evaluation strategies reported, and related both to programming experience, research area, and confidence ratings. Use was concentrated in five tasks: data handling, visualization, debugging, mathematical/scientific computing, and statistical analysis. Evaluation was informal and individual. Over half of accounts described running the generated code, while automated tests and review by another person were rare. Use cases and evaluation strategies varied little with programming experience, but confidence did: Less experienced programmers trusted the AI more than themselves, and experienced programmers the reverse. Evaluation confidence was not associated with the strategies reported. Its strongest correlates were confidence in the tool and in oneself. Validating AI contributions to scientific code rested largely on individual judgment, outside shared infrastructure for testing or review. Interfaces could support task-appropriate evaluation rather than leave it to the user.