基于多样性引导的搜索测试方法用于大型语言模型应用
Diversity-Guided Search-Based Testing of Large Language Model Applications
浏览论文内容
中文总结 AI 辅助
本文提出一种以失败多样性为优化目标的搜索测试框架,通过档案库奖励和重新填充操作符,在多个LLM应用上检测到更多失败并覆盖更广失败类型,其中贪婪重新填充权衡最佳。
中文摘要 AI 辅助
基于大型语言模型(LLM)的应用正越来越多地部署在客户服务、教育和出行等领域。这些系统容易产生不准确、虚构或有害的响应,其庞大且高维的输入空间使得系统性测试尤为困难。在本文中,我们提出了一种针对基于LLM的应用的基于搜索的测试框架,该框架将失败多样性作为明确的优化目标。基于风格、内容相关和扰动维度的离散化,我们的框架维护一个已生成测试的档案库,并奖励与档案库的距离,同时通过重新填充操作符定期替换种群中的非失败测试以维持探索。重新填充操作符由其采样策略参数化:替换候选要么均匀抽取,要么通过贪婪距离最大化抽取,我们实证比较了这两种设置。我们在三个案例研究——LLM安全性、车载导航和车内功能控制——中针对多个基线评估了这两种设置,覆盖了五个被测系统、八个LLM和18种不同的测试配置,执行了超过一百万次测试。我们的结果表明,在几乎所有配置中,所有引导变体检测到的失败数量都显著多于随机和组合搜索。其中,多样化搜索检测到的失败较少,但在所有案例研究中覆盖了更广泛的失败类型,而贪婪重新填充提供了最佳的权衡。
英文摘要
Large Language Model (LLM)-based applications are increasingly deployed across domains including customer service, education, and mobility. These systems are prone to inaccurate, fictitious, or harmful responses, and their vast, high-dimensional input space makes systematic testing particularly challenging. In this paper, we present a search-based testing framework for LLM-based applications that incorporates failure diversity as an explicit optimization objective. Building on a discretization along stylistic, content-related, and perturbation dimensions, our framework maintains an archive of generated tests and rewards distance from that archive, while a repopulation operator periodically replaces non-failing tests in the population to sustain exploration. The repopulation operator is parameterized by its sampling strategy: replacement candidates are drawn either uniformly or by greedy distance maximization, two settings we compare empirically. We evaluate both against several baselines across three case studies---LLM safety, in-car navigation, and in-vehicle function control---covering five systems under test, eight LLMs, and 18 distinct test configurations, with over one million executed tests. Our results show that all guided variants detect substantially more failures than random and combinatorial search in nearly all configurations. Among them, diversified search detects fewer failures, but covers a broader range of failure types in all case studies, with greedy repopulation offering the best tradeoff.
发表机构
- Technical University of Munich(慕尼黑工业大学)
- BMW Group(宝马集团)
- relAI – Konrad Zuse School for Reliable AI(relAI – 康拉德·楚泽可靠人工智能学院)
- fortiss GmbH(fortiss有限公司)
机构由 AI 辅助整理,请以论文原文为准。