arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关于通过行为代理实现模型代码与人类代码可理解性的行为对齐

On Behavioral Alignment of Model-Code and Human-Code Understandability via Behavioral Proxies

Xiaokai Rong, Aashish Yadavally, Anh H. N. Nguyen, Hridya Dhulipala, Tien N. Nguyen

arXiv 2609.26101首次发表:更新:

发表机构

University of Texas at Dallas; University of Central Florida(达拉斯得克萨斯大学; 中佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过行为代理量化模型代码可理解性,发现LLM与人类代码可理解性行为对齐优于基线,且与专业开发者对齐最紧密,语义自一致性是可靠度量。

AI 中文摘要

代码可理解性是软件质量的一个关键方面。以往的研究主要从以人为中心或代码为中心的角度来关注这一属性,而它应被视为读者与代码之间交互产生的关系属性。随着大语言模型在软件工程中的日益普及,我们认为“读者”的概念应被泛化,以涵盖人类和模型。基于这种关系视角,我们扩展了代码可理解性的概念,以区分人类代码可理解性和模型代码可理解性,旨在研究两者之间的行为对齐。为此,我们使用了先前研究中的数据集,该数据集包含不同参与者群体对人类可理解性的评级判断,并评估了多个开源和闭源LLM。为了实现这种行为对齐的操作化,我们引入了四个模型代码可理解性的行为代理(BPMU):P0,基于程序理解聚焦的问答;以及P1--P3,基于程序意图摘要(源自我们提出的语义自一致性)。我们的研究结果表明,LLM在行为对齐方面与一般人类代码可理解性的表现优于先前使用的浅层机器学习基线。分层分析进一步揭示,模型代码可理解性与专业开发者的对齐最为紧密,但与其他学生群体的对齐显著较少。角色条件提示并未改善后者的对齐,这表明当前的LLM无法准确模仿具有不同专业水平的人类读者的视角。最后,我们证明了语义自一致性是一种可靠且可扩展的度量,可作为量化模型代码可理解性的行为代理,对软件工程研究和实践具有广泛意义。

英文摘要

Code understandability is a critical aspect of software quality. Prior research has largely focused on this attribute from a human-centric or code-centric perspective, while it should be viewed as a relational property arising from the interaction between a reader and the code. With the increasing adoption of large language models in software engineering, we posit that the notion of "reader" should be generalized to encompass both humans and models. Building on this relational perspective, we extend the concept of code understandability to distinguish between human and model code understandability, aiming to investigate the behavioral alignment between the two. To this end, we use a dataset from a prior study containing human-rated judgments of understandability across diverse participant groups, and evaluate multiple open and closed-source LLMs. To operationalize such behavioral alignment, we introduce four behavioral proxies of model-code understandability (BPMU): P0, based on program comprehension-focused question answering; and P1--P3, based on program intent summarization (derived from our proposed semantic self-consistency). Our findings show that LLMs exhibit stronger behavioral alignment with general human code understandability than previously used shallow machine learning baselines. Stratified analyses further reveal that model code understandability aligns most closely with professional developers, but significantly less with other student groups. Role-conditioned prompting does not improve the alignment for the latter, suggesting that current LLMs cannot accurately mimic the perspectives of human readers with varying expertise. Finally, we demonstrate that semantic self-consistency is a reliable and extensible measure to be used as a behavioral proxy for quantifying model code understandability, with broad implications in both software engineering research and practice.

DOI:10.1145/3832202

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑