arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.28775cs.SE

基于大型语言模型的增强现实应用的仓库感知变形关系生成

Repository-Aware Metamorphic Relation Generation for Augmented Reality Applications using Large Language Models

Dibyendu Brinto Bose, Jiawei Qin, Chris Brown

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种结合仓库级上下文与推理编排的流水线,利用大型语言模型生成优化的变形关系,在142个移动AR仓库数据集上验证了该方法可构建可靠的领域相关测试预言,能有效检测代码变异。

中文摘要 AI 辅助

变形测试(Metamorphic Testing, MT)提供了一种无需定义测试预言即可测试软件的有前景方法,该方法通过指定输入与输出之间的预期关系而非依赖精确输出来实现。例如,测试增强现实(Augmented Reality, AR)应用极具挑战性,因为虚拟内容、物理环境与代码之间存在动态交互,这使得传统测试预言难以定义。然而,构建变形关系(Metamorphic Relations, MRs)既耗时又繁琐。本文提出一种上下文感知流水线,利用仓库级上下文与推理编排来生成并优化MRs,在包含142个移动AR系统仓库的数据集上进行评估。在生成14916个候选MRs的三种上下文配置中,分层上下文产生最广泛的覆盖范围(142个仓库中共7004个MRs,涉及5167个类-方法对)且冗余度更低。随后,智能体审议过程协调冲突候选(79.0%的案例中存在此类冲突),减少重复并在88.2%的结果中选择上下文感知关系。人工预言研究显示,优化后的关系(n=141)在逻辑上有效且足够具体,可直接转换为测试断言;初步案例研究表明,将生成的MRs(n=5)转换为可执行测试能够检测真实代码中的非等价变异。总体而言,我们的结果表明,将仓库感知的MR生成与基于推理的优化相结合,可实现可靠、领域相关的测试预言的可扩展构建。

英文摘要

Metamorphic Testing (MT) provides a promising approach for testing software without defined test oracles by specifying expected relations between inputs and outputs, instead of relying on exact outputs. For example, testing Augmented Reality (AR) applications is challenging due to dynamic interactions between virtual content, physical environments, and code, which make traditional test oracles difficult to define. However, formulating metamorphic relations (MRs) is time-consuming and burdensome. We introduce a context-aware pipeline that generates and refines MRs using repository-level context and reasoning orchestration, evaluated on a dataset of 142 mobile AR system repositories. Across three context configurations generating 14,916 candidate MRs, hierarchical context yielded the broadest coverage (7,004 MRs across 142 repositories and 5,167 class--method pairs) and lower redundancy. An agentic deliberation process then reconciled conflicting candidates---observed in 79.0% of cases---reducing duplication and selecting context-aware relations in 88.2% of outcomes. A manual oracle study shows refined relations (n = 141) are both logically valid and sufficiently concrete to be directly translated into test assertions, and a preliminary case study reveals converting generated MRs (n = 5) into executable tests can detect non-equivalent mutations in real-world code. Overall, our results show that combining repository-aware MR generation with reasoning-based refinement enables scalable construction of reliable, domain-relevant test oracles.

↑