发表机构
University of Oviedo; Munster Technological University(奥维耶多大学; 芒斯特理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SafeLLM4SE提出一种统计原则性的评估与报告方法,将LLM输出视为随机过程实现,结合自适应采样、置信区间和效应量,并开源PyPI包,在HumanEval上演示了应用。
AI 中文摘要
大型语言模型(LLM)越来越多地用于软件工程任务,然而其随机行为对评估的有效性、可复现性和可比性构成了挑战。诸如报告单一输出、平均分数、最佳N(best-of-N)或pass@k性能等传统做法可能会掩盖变异性和估计不确定性,从而可能对系统可靠性得出误导性结论。本文提出了SafeLLM4SE,一种用于对基于LLM的软件工程系统进行统计原则性评估的实用方法论和报告标准。SafeLLM4SE并未将生成的输出视为确定性产物,而是将其视为随机过程的实现,并区分质量、稳定性和估计不确定性。它结合了自适应采样与置信区间、考虑分布差异的统计比较、效应量,以及涵盖模型配置、可复现性、评估程序和资源使用的最低报告标准。SafeLLM4SE还作为开源软件包在PyPI上提供,使研究人员和从业者能够复现和扩展该方法。我们通过比较两个LLM在HumanEval(一个通过功能测试评估的编程问题基准)上的表现来展示其应用。
英文摘要
Large language models (LLMs) are increasingly used for software engineering tasks, yet their stochastic behavior challenges the validity, reproducibility, and comparability of their evaluations. Conventional practices such as reporting a single output, an average score, best-of-N, or pass@k performance can obscure variability and estimation uncertainty, potentially leading to misleading conclusions about system reliability. This article presents SafeLLM4SE, a practical methodology and reporting standard for statistically principled evaluation of LLM-based software engineering systems. Rather than treating generated outputs as deterministic artifacts, SafeLLM4SE treats them as realizations of a stochastic process and distinguishes quality, stability, and estimation uncertainty. It combines adaptive sampling with confidence intervals, distribution-aware statistical comparisons, effect sizes, and a minimum reporting standard covering model configuration, reproducibility, evaluation procedures, and resource usage. SafeLLM4SE is also provided as an open-source software package available on PyPI, enabling researchers and practitioners to reproduce and extend the methodology. We illustrate its application by comparing two LLMs on HumanEval, a benchmark of programming problems assessed through functional tests.
CommentsThis work has been submitted for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible