发表机构
Independent Researcher, USA; Technical University of Munich, Germany(美国独立研究员; 德国技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
通过系统文献综述,沿权威来源、形式及裁决机制三轴分析相关研究,刻画领域、语言等情况,指出规范衍生权威最常见但仅占约一半研究,还报告了评估及失败情况并提出研究议程。
AI 中文摘要
大语言模型(LLMs)越来越多地用于生成测试预言机,但其权威来源尚不明确。以往研究未按裁决权威来源分类。本文按PRISMA 2020指南进行系统文献综述,分析了54项研究,刻画了相关情况,指出规范衍生权威虽常见但占比一半,还报告了评估及失败情况并将分类法空白处作为研究议程。
英文摘要
Large language models (LLMs) increasingly decide whether software behaves correctly, either by writing a test oracle or by acting as one. Yet two oracles can look identical and rest on different ground: one assertion encodes a written specification, another only what the model learned in training. Prior secondary studies sort oracles by form or by technique, rarely by the property that governs how far a verdict can be trusted: where its authority comes from. This systematic literature review, reported under the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines, screens 2,436 records to 54 included studies, extended by citation searching (snowballing) to 83 in total. We read the corpus along three axes: the source of an oracle's authority, the form it takes, and the mechanism that adjudicates it. Just over half of the corpus reaches a verdict with no specification at all. That is what lets these oracles work on code with no specification to consult, and what leaves a challenged verdict with less to fall back on. Source and mechanism cross-cut rather than coincide, so a label such as LLM-as-a-judge names how a verdict is produced, not why it should be trusted. Oracle quality is most often judged by resemblance to a known oracle rather than by whether injected faults are caught. The first question to ask of any LLM oracle is therefore what one would point to in defending its verdict. The protocol, search query, and per-study coding sheet are released.
Comments21 pages, 10 figures, 11 tables. Systematic literature review of 83 studies, reported under PRISMA 2020. Published in IEEE Access. Replication package: https://doi.org/10.5281/zenodo.21194940
Journal refIEEE Access, vol. 14, 2026
DOI:10.1109/ACCESS.2026.3729738