arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Code-MUE:通过基于执行的语义交互图测量代码语言模型的不确定性

Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs

Xiaoning Ren, Yinxing Xue, Lei Ma, Yuheng Huang

arXiv 2607.12273首次发表:更新:

发表机构

Xi’an Jiaotong University; Institute of AI for Industries, Chinese Academy of Sciences; The University of Tokyo; University of Alberta(西安交通大学; 中国科学院人工智能产业研究院; 东京大学; 阿尔伯塔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对代码语言模型内在随机性带来的风险,引入纯黑盒框架Code-MUE,通过基于执行的语义交互图测量不确定性,经大规模实证研究验证其与功能正确性强负相关,优于基线,可实现风险检测和选择性预测。

AI 中文摘要

随着代码大语言模型(LLMs)成为现代软件工程的核心,其内在的随机性带来了重大现实风险,小错误也可能导致严重后果。现有不确定性估计方法存在差距,白盒和灰盒技术不适用于闭源模型,标准黑盒文本指标无法捕捉代码独特脆弱性。为此引入Code-MUE,一个通过基于执行的语义交互图测量不确定性的纯黑盒框架。它基于可观察运行时行为计算解空间的冯·诺依曼熵来量化全局语义多样性。大规模实证研究表明,Code-MUE与功能正确性呈强负相关,显著优于基于词汇和嵌入的基线,能在实际工作流程中实现强大的风险检测和选择性预测。

英文摘要

As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, where even minor errors can lead to severe functional, security, or safety consequences. Reliable automation, therefore, demands the ability to distinguish between confident, well-supported predictions and stochastic guessing. However, existing uncertainty estimation methods face a critical gap: white and grey-box techniques are often inapplicable to closed-source models, while standard "black-box" text metrics fail to capture the unique fragility of code, where syntactic variation does not always imply semantic divergence. To bridge this syntax-semantics gap, we introduce Code-MUE, a purely black-box framework that measures uncertainty through execution-based Semantic Interaction Graphs. Different from prior approaches that rely on superficial textual similarity, Code-MUE grounds uncertainty in observable runtime behavior, calculating the Von Neumann entropy of the solution space to quantify global semantic diversity. A large-scale empirical study across eight state-of-the-art LLMs demonstrates that Code-MUE achieves a strong negative correlation with functional correctness (Spearman's correlation up to -0.98), significantly outperforming lexical and embedding-based baselines while enabling robust risk detection and selective prediction in practical workflows.

CommentsTo appear at The ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑