AI 中文总结
本文通过在三个数据集、四种模型架构下的20种配置开展实证研究,发现解释引导式变异测试生成的经验证故障诱导测试用例是启发式变异策略的2.30倍,可有效评估专用语言模型的鲁棒性。
AI 中文摘要
背景:任务专用语言模型正越来越多地被集成到软件工程工作流中,以支持问题分类、文档分类和自动化分析等垂直领域活动。尽管它们得到了应用,但关于如何在语义保留输入变换下测试其鲁棒性并检测脆弱行为的实证证据有限。目标:本文研究与启发式变异策略相比,解释引导式变异测试是否能提升专用语言模型鲁棒性测试的有效性和合理性。方法:我们在三个数据集、四种模型架构及20种由归因方法与变异策略组合而成的测试配置下,开展了对解释引导式变异测试的大规模实证研究。所评估的配置结合了基于归因的词元优先级排序、大语言模型(LLM)驱动的变异及自动化语义验证,以生成语言有效的测试变体。我们针对启发式基线评估了故障发现能力、语义有效性及测试效率。结果:解释引导式变异测试生成的经验证的故障诱导测试用例是启发式变异策略的2.30倍。语义验证大幅提升了变异有效性,并在人工标注的门控接受变体中实现了高标签保留精度。该研究进一步揭示了模型中存在的系统性捷径行为,包括过度依赖命名实体和格式线索。结论:结果提供证据表明,解释引导式变异测试是一种有效且实用的方法,可用于对垂直AI应用中使用的任务专用语言模型的鲁棒性进行实证评估。
英文摘要
\head{Background} Task-specialized language models are increasingly integrated into software engineering workflows to support vertical-domain activities such as issue triaging, document classification, and automated analysis. Despite their adoption, there is limited empirical evidence on how to test their robustness and detect brittle behaviors under semantics-preserving input transformations. \head{Aims} This paper investigates whether explainability-guided metamorphic testing can improve the effectiveness and validity of robustness testing for specialized language models compared to heuristic mutation strategies. \head{Method} We conduct a large-scale empirical study of explanation-guided metamorphic testing across three datasets, four model architectures, and 20 testing configurations derived from combinations of attribution methods and mutation strategies. The evaluated configurations combine attribution-based token prioritization, LLM-driven mutation, and automated semantic verification to generate linguistically valid test variants. We assess failure discovery capability, semantic validity, and testing efficiency against heuristic baselines. \head{Results} Explanation-guided metamorphic testing generates 2.30$\times$ more verified failure-inducing test cases than heuristic mutation strategies. Semantic verification substantially improves mutation validity and achieves high label-preservation precision among gate-accepted variants according to human annotation. The study further reveals systematic shortcut behaviors across models, including over-reliance on named entities and formatting cues. \head{Conclusions} The results provide evidence that explanation-guided metamorphic testing is an effective and practical approach for empirically evaluating the robustness of task-specialized language models used in vertical AI applications.
Comments20 pages, 4 figures. Accepted at ESEM 2026