arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04195cs.SE

SONAR:面向LLM使用者的任务感知型代码摘要评估框架(无需参考摘要)

SONAR: Task-Aware Code Summary Evaluation for LLM Consumers Without References

Simantika Bhattacharjee Dristi, Matthew B. Dwyer

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出无需参考摘要的SONAR框架,从四个维度评估代码摘要,发现正确性、抽象性对LLM性能影响显著,简洁性、流畅性影响微弱,还明确了不同LLM在各维度的优劣。

中文摘要 AI 辅助

传统上,源代码摘要的评估从人类开发者的视角出发,质量取决于其与开发者编写的参考摘要的相似程度,以及与人类偏好的契合度。但这忽略了一个日益凸显的现实:基于大语言模型(LLM)的工具和智能体越来越多地将代码摘要作为软件工程(SE)任务的输入,而摘要对执行任务的消费型智能体而言具备实用性的核心因素,在很大程度上仍未被探究。为填补这一空白,我们提出SONAR,这是一个无需参考摘要的框架,从四个维度评估源代码摘要:正确性(Correctness)、抽象性(Abstraction)、简洁性(Conciseness)和流畅性(Fluency)。SONAR不针对预先编写的“黄金标准”进行优化,而是引入了一种新颖的基于代码再生的方法:利用摘要再生代码,并将这种重构过程作为摘要质量的信号,这提供了既无需参考摘要、也无需人类或LLM主观判断的经验依据。我们在四个下游SE任务中评估了SONAR各维度对LLM性能的影响能力,发现正确性(其次是抽象性)与LLM性能显著相关,其相关性比最佳基准方法高出14倍。尽管简洁性和流畅性受到人类开发者的广泛重视,但对LLM消费者而言大多无显著影响,这表明摘要的实用性取决于任务和消费对象。通过使用SONAR对11种流行LLM进行大规模评估,我们进一步明确了不同模型在各质量维度上的优势与劣势,同时为未来任务感知型摘要研究提供了见解。

英文摘要

Source code summaries have traditionally been evaluated from a human developer's perspective, with quality determined by how closely they resemble developer-written references and how well they align with human preferences. But this overlooks a growing reality: LLM-based tools and agents increasingly consume code summaries as inputs for software engineering (SE) tasks, and what makes a summary useful for a consuming agent on a task remains largely unexplored. To bridge this gap, we propose SONAR, a reference-free framework that evaluates source code summaries along four dimensions: Correctness, Abstraction, Conciseness, and Fluency. Rather than optimizing for a pre-written "gold standard", SONAR introduces a novel code regeneration-based approach that uses a summary to regenerate code and leverages that reconstruction as a quality signal of the summary. This provides an empirical grounding that requires neither a reference summary nor the subjective judgment of humans or LLMs. We evaluate SONAR's dimensions on their ability to influence LLM performance across four downstream SE tasks. We find that Correctness, followed by Abstraction, significantly correlates with LLM performance, with correlations up to 14X higher than the best baseline. Conciseness and Fluency, though widely valued by human developers, remain mostly insignificant to an LLM consumer, suggesting that what makes a summary useful is task- and consumer-dependent. Through a large-scale evaluation of 11 popular LLMs using SONAR, we further identify the strengths and weaknesses of different models across each quality dimension, while offering insights to facilitate future research on task-aware summarization.

↑