发表机构
Xi’an Jiaotong University; Institute of AI for Industries, Chinese Academy of Sciences; The University of Tokyo; University of Alberta(西安交通大学; 中国科学院人工智能产业研究院; 东京大学; 阿尔伯塔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对代码语言模型内在随机性带来的风险,引入纯黑盒框架Code-MUE,通过基于执行的语义交互图测量不确定性,经大规模实证研究验证其与功能正确性强负相关,优于基线,可实现风险检测和选择性预测。
AI 中文摘要
随着代码大语言模型(LLMs)成为现代软件工程的核心,其内在的随机性带来了重大现实风险,小错误也可能导致严重后果。现有不确定性估计方法存在差距,白盒和灰盒技术不适用于闭源模型,标准黑盒文本指标无法捕捉代码独特脆弱性。为此引入Code-MUE,一个通过基于执行的语义交互图测量不确定性的纯黑盒框架。它基于可观察运行时行为计算解空间的冯·诺依曼熵来量化全局语义多样性。大规模实证研究表明,Code-MUE与功能正确性呈强负相关,显著优于基于词汇和嵌入的基线,能在实际工作流程中实现强大的风险检测和选择性预测。
英文摘要
As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, where even minor errors can lead to severe functional, security, or safety consequences. Reliable automation, therefore, demands the ability to distinguish between confident, well-supported predictions and stochastic guessing. However, existing uncertainty estimation methods face a critical gap: white and grey-box techniques are often inapplicable to closed-source models, while standard "black-box" text metrics fail to capture the unique fragility of code, where syntactic variation does not always imply semantic divergence. To bridge this syntax-semantics gap, we introduce Code-MUE, a purely black-box framework that measures uncertainty through execution-based Semantic Interaction Graphs. Different from prior approaches that rely on superficial textual similarity, Code-MUE grounds uncertainty in observable runtime behavior, calculating the Von Neumann entropy of the solution space to quantify global semantic diversity. A large-scale empirical study across eight state-of-the-art LLMs demonstrates that Code-MUE achieves a strong negative correlation with functional correctness (Spearman's correlation up to -0.98), significantly outperforming lexical and embedding-based baselines while enabling robust risk detection and selective prediction in practical workflows.
CommentsTo appear at The ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA) 2026