AI 中文总结
该研究针对基于LLM的定向测试输入生成技术的不足,提出ReDig框架,通过运行时反馈的控制循环优化测试输入生成,在Poppler和Libsndfile案例中验证了其诊断失败原因与优化测试脚本的有效性。
AI 中文摘要
基于大语言模型(LLM)的定向输入生成技术在生成可到达目标的测试输入方面已展现出良好效果。然而,由于可用代码信息的限制以及LLM推理的固有不可预测性,可靠的定向输入生成需要将该过程基于观测到的运行时行为。我们提出ReDig,这是一个由运行时反馈引导的优化框架,它在基于LLM的定向测试输入生成技术周围添加了一个控制循环,以利用之前未到达目标的测试执行中观测到的运行时值来优化定向输入生成。在Poppler和Libsndfile的案例研究中,我们发现ReDig能有效获取运行时值反馈,以诊断之前的测试脚本未能到达目标行的原因,还能有效利用给定的运行时值反馈在后续步骤中优化测试脚本。
英文摘要
LLM-based directed input generation techniques have shown promising effectiveness at producing target-reaching test inputs. However, due to the constraint of available code information and inherent unpredictability of LLM inference, reliable directed input generation requires mechanisms to ground the process in observed runtime behavior. We propose ReDig, a runtime feedback-guided refinement framework which adds a control loop around an LLM-based directed test input generation technique to refine the directed input generation with runtime values observed in prior target-missing test executions. In the case studies with Poppler and Libsndfile, we found that ReDig effectively derive runtime value feedback to diagnose why the previous test script failed to reach the target lines, and also effectively leverage given runtime value feedback to refine the test scripts in subsequent steps.