arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

开发者使用本地大语言模型生成和评估的代码导航体验陌生代码库调试的研究

How Developers Experience Debugging Unfamiliar Codebases with Code Tours Generated and Evaluated by Local LLMs

Martin Balfroid, Julien Albert, Dzenatan Aliti, Xavier Devroey, Benoît Vanderose

arXiv 2607.26987首次发表:更新:

AI 中文总结

该研究探讨开发者使用本地大语言模型生成评估的代码导航调试陌生代码库的体验,明确代码导航组件属性对开发者的影响,为优化代码导航生成与评估提供方向。

AI 中文摘要

代码导航是一种交互式的入门文档,用于引导开发者浏览代码库。大语言模型(LLM)可自动合成代码导航。此前关于代码导航生成的研究,尚未探讨开发者使用开源权重LLM生成和评估的代码导航调试陌生代码库时的体验或信任校准问题。本研究调查了开源权重LLM生成的代码导航中组件属性如何影响开发者调试陌生代码库时的体验。我们构建了一个从真实可复现bug生成和评估代码导航的流程。共有26名不同背景的开发者参与了用户研究。总共从2025年GitHub提交中挖掘的真实Java bug生成了26个代码导航,每个导航由两个不同的LLM独立评判,形成52种评估配置。参与者在探索每个导航时进行出声思考。三名作者对访谈进行定性编码以识别反复出现的主题。开发者通常倾向于那些细节随代码长度扩展、避免仅复述代码、易于扫描且采用引导语气的导航。然而,一些偏好相互排斥,例如祈使语气的使用。栈跟踪通常不足以识别开发者认为相关的所有步骤。开发者信任他们认为是人工编写的描述,而非他们认为是AI生成的描述。最后,LLM生成的导航质量标注不可靠:奉承、虚构和不连贯现象普遍存在。本研究为未来研究奠定了基础,包括对开源权重模型进行代码导航生成的微调、个性化生成以适应不同偏好、选择栈跟踪之外的相关步骤、校准用户信任以避免闲置和误用,以及提高开源权重LLM作为更可信评估者的能力。

英文摘要

Code tours are interactive, onboarding documentation to guide developers through a codebase. Large Language Models (LLMs) can automatically synthesize code tours. Prior work on code tour generation has not studied developer experience or trust calibration when debugging unfamiliar codebases with code tours generated and evaluated by open-weight LLMs. This study surveys how the properties of components in open-weight LLM-authored code tours influence developers' experiences when debugging unfamiliar codebases. We built a pipeline that generated and evaluated code tours from real reproducible bugs. 26 developers with varying backgrounds participated in a user study. In total, 26 code tours were authored from real Java bugs mined from 2025 GitHub commits, with each tour independently judged by two different LLMs, resulting in 52 evaluated configurations. Participants thought aloud as they explored each tour. Three authors qualitatively coded the interviews to identify recurring themes. Developers generally preferred tours that scaled detail with the code length, avoided merely restating code, were easily scannable, and adopted a guiding tone. However, some preferences were mutually exclusive, such as the use of imperative mood. Stack traces were often insufficient to identify all steps developers found relevant. Developers also trusted descriptions they perceived as human-written more than those they believed were AI-generated. Finally, LLM-generated annotations of tour quality were unreliable: sycophancy, confabulation, and incoherence were pervasive. This work lays a basis for future research on fine-tuning open-weight models for code tour generation, personalizing generation to accommodate diverging preferences, selecting relevant steps beyond stack traces, calibrating users' trust to avoid both disuse and misuse, and improving open-weight LLMs' ability to be more trustworthy evaluators

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑