arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于智能体的多视角聚合测试断言生成

Agent-Based Test Assertion Generation via Diverse Perspective Aggregation

Dong Wang, Qiaoyu Han, Lin Yang, Jianyi Zhou, Guangtai Liang, Junjie Chen

arXiv 2608.05822首次发表:更新:

AI 中文总结

针对现有LLM断言生成方法的局限,提出基于智能体的AssertMate框架,通过三个组件聚合多视角,在Defects4J和EvoSuite验证中性能显著优于现有技术。

AI 中文摘要

测试断言是单元测试的关键元素,作为检查点用于验证预期行为并确保软件正确性。已有诸多技术被提出以实现断言生成自动化,近期进展尤其由大语言模型(LLM)推动。尽管有前景,现有方法如ChatAssert存在准确率偏低、严重依赖过采样、因单次提示易受模型随机性影响的问题。为解决这些局限,我们提出AssertMate,一种新型基于智能体的断言生成框架,通过三个关键组件提升LLM生成断言的质量与可靠性:(1)实际值构建,通过静态分析和感知类型的启发式方法识别断言目标;(2)多视角预期值预测,采用代码生成、检索增强生成(RAG)及思维链(CoT)推理智能体;(3)LLM作为评判者的协作机制以选择最恰当的断言。在Defects4J基准上的评估表明,AssertMate在编译成功率和通过率上显著优于现有技术,且具有高得多的缺陷检测能力;与EvoSuite的集成进一步验证了其实用性,产生更优的变异覆盖率和变异杀死数。消融研究显示三个组件均对整体性能有显著且互补的贡献。本研究证实聚合多视角可提升基于LLM的断言生成的有效性。

英文摘要

Test assertions are critical elements of unit tests, serving as checkpoints to validate expected behavior and ensure software correctness. Numerous techniques have been proposed to automate assertion generation, with recent progress notably driven by large language models (LLMs). Despite the promise, existing approaches such as ChatAssert suffer from modest accuracy, heavy reliance on oversampling, and vulnerability to model randomness due to one-shot prompting. To address these limitations, we propose AssertMate, a novel agent-based assertion generation framework that enhances the quality and reliability of LLM-generated assertions through three key components: (1) actual value construction that identifies assertion targets via static analysis and type-aware heuristics; (2) multi-perspective expected value prediction using code generation, retrieval-augmented generation (RAG), and chain-of-thought (CoT) reasoning agents; and (3) an LLM-as-a-Judge collaboration mechanism to select the most appropriate assertion. Evaluation on the Defects4J benchmark demonstrates that AssertMate significantly outperforms state-of-the-art techniques in compilation success and pass rates, along with substantially higher bug detection capabilities. Integration with EvoSuite further validates AssertMate's practicality, yielding superior mutation coverage and kill counts. Ablation studies reveal that each of the three components makes a significant and complementary contribution to the overall performance. This work affirms the great potential of aggregating diverse perspectives to enhance the effectiveness of LLM-based assertion generation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑