AI 中文总结
该研究针对基于大语言模型的软件的验收测试缺口,提出需求增强生成技术与置信度校准级联判定方法,经工业案例验证可提升预言机质量、准确率与成本效率,具备工业可行性。
AI 中文摘要
基于大语言模型的软件(LBS)将大语言模型作为核心组件,以提供灵活、个性化的响应。与具有确定性输出的传统软件不同,LBS表现出依赖上下文的随机行为,这使得经典验收测试和测试预言机变得不足:同一查询可能根据用户角色和软件上下文需要完全不同的响应。这一缺口催生了对自动化验收测试框架的迫切需求,该框架能自主解读用户指令,同时在变化的环境中可靠推断用户意图。本文提出一种面向LBS的自动化验收测试框架,通过两项技术贡献实现校准后的判定可靠性。首先,我们引入需求增强生成技术(REAG),该技术通过自适应检索增强生成(RAG)和自我推理检索相关的软件需求、领域知识和用户角色,以解读用户意图并生成上下文感知的测试预言机。其次,考虑到预言机生成可能检索到不相关的约束、误读意图或生成幻觉需求,我们引入置信度校准级联判定方法。该方法通过模拟专家一致性来量化判定可靠性:接受高置信度判定、升级模糊案例,或在不确定时弃权(不执行),并通过保形风险控制提供经验可靠性保证。针对工业级营养咨询应用的案例研究表明,REAG的预言机质量得分为3.91/5,在82%的案例中达到合格或边缘级预言机质量;置信度校准级联判定的准确率达98.8%,通过过滤不合格输出将预言机质量从3.91提升至4.30,相比单一判定基准实现了31.7%的成本效率提升,验证了其工业可行性。
英文摘要
LLM-based software (LBS) integrates large language models as core components to deliver flexible, personalised responses. Unlike traditional software with deterministic outputs, LBSs exhibit context-dependent, stochastic behaviour that renders classical acceptance testing and test oracles insufficient: the same query may require fundamentally different responses depending on user personas and software context. This gap creates an urgent need for automated acceptance testing frameworks that autonomously interpret user instructions, while reliably inferring user intentions in a changing environment. In this paper, we present an automated acceptance testing framework for LBS with calibrated verdict reliability via two technical contributions. First, we introduce Requirements-Augmented Generation (REAG), which interprets user intentions by retrieving relevant software requirements, domain knowledge, and personas via adaptive RAG and self-reasoning to generate context-aware test oracles. Second, recognising that oracle generation may retrieve irrelevant constraints, misinterpret intent, or hallucinate requirements, we introduce a confidence-calibrated cascade judgment. This method quantifies verdict reliability via simulated expert agreement -- accepting high-confidence verdicts, escalating ambiguous cases, or abstaining when uncertain -- with empirical reliability guarantees backed by conformal risk control. An industrial case study on a production nutrition advisory application demonstrates that REAG achieves a 3.91/5 oracle quality score, reaching qualified or marginal oracle quality in 82% of cases. The confidence-calibrated cascade achieves 98.8% accuracy, improves oracle quality from 3.91 to 4.30 by filtering unqualified outputs, and delivers a 31.7% cost-efficiency improvement over single-judge baselines, validating industrial viability
CommentsAccepted at ASE2026