arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

回到未来:从LLM回归可读性特征

Back to the Future: Regressing Readability Features from LLMs

Tairan Wang, Earl T. Barr

arXiv 2610.04641首次发表:更新:

发表机构

University College London(伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有可读性模型与人类判断不一致的问题,提出BTTF方法,将线性回归与信息论特征结合,利用LLM作为测量仪器,在五个基准和1,100个样本上显著提升Spearman相关性,并具备特征级可解释性。

AI 中文摘要

代码可读性支持软件理解、审查和维护。随着AI编码智能体被更广泛地使用,可读的代码既有助于人类审查智能体生成的输出,也有助于智能体在有限的上下文窗口内维护代码。然而,现有的可读性模型在不同数据集上与人类判断的一致性并不稳定。我们引入了BTTF(回到未来),将线性回归与信息论特征相结合。BTTF不是将语言模型用作黑盒评判者,而是将其用作测量仪器。它结合了传统代码度量、基于嵌入的语义组织以及因果语言模型的可预测性,以捕获源代码结构、语义连贯性和上下文意外性。我们在五个公开的人工评级可读性基准和我们手动标注的MBJP开发集上评估了BTTF,总计1,100个样本。BTTF在平均Spearman相关性上比最强基线提高了最多0.101。在对生产代码进行的13种受控可读性降低转换中,它在Java中实现了93.8%的响应率,在Python中实现了87.6%的响应率,分别优于Dorn 14%和19.2%。它在本地模型中表现最佳,在Java中落后最强云LLM响应1.7%,在Python中落后5.1%。直接LLM基线每次评估平均需要1,083个提示词token,且保持不透明。BTTF结合了最先进的基准性能和强大的转换鲁棒性,并具有特征粒度的可解释性:每个特征系数在所有评估的模型配置中都与其理论方向效应匹配。

英文摘要

Code readability supports software understanding, review, and maintenance. As AI coding agents become more widely used, readable code helps both humans review agent-generated output and agents maintain code within limited context windows. Yet existing readability models do not consistently agree with human judgements across datasets. We introduce BTTF (Back To The Future), pairing linear regression with information-theoretic features. Rather than using language models as black-box judges, BTTF uses them as measurement instruments. It combines traditional code measurements, embedding-based semantic organisation, and causal-LM predictability to capture source structure, semantic coherence, and contextual surprise. We evaluate BTTF on five publicly available human-rated readability benchmarks and our manually annotated MBJP development set, totalling 1,100 samples. BTTF improves average Spearman correlation over the strongest baseline by up to 0.101. Across 13 controlled readability-reducing transformations on production code, it achieves response rates of 93.8% in Java and 87.6% in Python, outperforming Dorn by 14% and 19.2%. It performs best among local models and trails the strongest cloud-LLM response by 1.7% in Java and 5.1% in Python. The direct-LLM baselines require an average of 1,083 prompt tokens per evaluation while remaining opaque. BTTF combines state-of-the-art benchmark performance and strong transformation robustness with feature-granular interpretability: every feature coefficient matches its theoretical directional effect across all evaluated model configurations.

Comments39 pages, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑