arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用模型检验作为神谕对基于大语言模型(LLM)的事后解释器进行自动化测试

Automated Testing of LLM-Based Post Hoc Explainers Using Model Checking as an Oracle

Dennis Gross, Helge Spieker

arXiv 2608.30581首次发表:更新:

发表机构

Institut für Kommunikations- und Prüfungsforschung gGmbH; Simula Research Laboratory(通信与测试研究院(非营利有限责任公司); Simula研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出以概率模型检验为神谕、结合事后查询分类法,实现对基于LLM的事后解释器的自动化测试,在7个MDP环境中区分出不同规模LLM的解释可信度。

AI 中文摘要

大语言模型(LLM)被用作序贯决策策略的事后解释器,生成关于为何选择某一动作的自然语言解释。然而,LLM常生成看似合理实则错误的表述,且现有方法均未系统测试此类解释是否忠实于底层环境。两大经典软件测试挑战阻碍了相关工作:一是不存在解释正确性的神谕,二是测试输入(关于策略行为的自然语言查询)缺乏系统生成测试用例所需的结构。我们解决了这两个问题:概率模型检验提供了测试神谕,可计算精确参考结果以自动评判LLM的回答;事后查询类别的分类法围绕策略解释所组成的环境级事实构建输入空间,据此生成的测试用例按特定问题的诊断难度评分排序。在7个马尔可夫决策过程(MDP)环境中,该测试对3个开放权重LLM进行了区分:一个推理模型通过了85%的测试用例,一个中等规模模型通过了70%,而一个10亿参数的模型低于随机基线;同时,难度排序机制生成的测试用例比随机选择的更具挑战性。我们的结果揭示了在无模型设置中LLM生成解释的可信度,而在这类设置中,虽使用相同LLM,但不存在可验证它们的神谕。

英文摘要

Large language models (LLMs) are used as post hoc explainers of sequential decision-making policies, producing natural-language explanations of why an action was chosen. However, LLMs often generate plausible but incorrect statements, and no existing approach systematically tests whether such explanations are faithful to the underlying environment. Two classic software testing challenges stand in the way: there is no oracle for the correctness of an explanation, and the test inputs, natural language queries about a policy's behavior, lack the structure needed for systematic test case generation. We address both. Probabilistic model checking provides the test oracle, computing exact reference results against which LLM answers are graded automatically. A taxonomy of post hoc query categories structures the input space around the environment-level facts from which policy explanations are composed; test cases generated from it are prioritized by question-specific diagnostic difficulty scores. Across seven MDP environments, the testing separates three open-weight LLMs: a reasoning model passes 85% of test cases, a mid-size model 70%, and a 1B model falls below the random baseline, while prioritization surfaces significantly harder cases than random selection. Our results indicate how trustworthy LLM-generated explanations are in model-free settings, where the same LLMs are used but no oracle exists to verify them.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑