基于句子级能量景观的黑盒大语言模型解释
Interpreting Black-Box Large Language Models with Sentence-Level Energy Landscapes
浏览论文内容
中文总结 AI 辅助
针对黑盒LLM的可解释性缺失问题,提出一种句子级模型无关后验归因解释器,通过EBM代理捕获概念一致性,无需额外API查询即可量化提示对目标输出的影响,实验验证其能有效识别关键提示。
中文摘要 AI 辅助
通过封闭API严格访问的专有大语言模型(LLM)的广泛应用,给负责任部署带来了关键挑战:根本性的可解释性缺失。为解决这一问题,我们提出一种模型无关的后验归因解释器,在句子级别运行。该方法训练一个基于能量的模型(EBM)作为代理,以捕获LLM在提示与响应之间的内部概念一致性;此能量景观指导轻量解释器网络的训练。我们的解释器独特之处在于作为独立工具运行:一旦训练完成,无需再向LLM查询API,即可量化提示句子对用户指定目标输出的影响。通过在多样化输入上全局训练局部解释器,我们的框架捕捉更广泛的生成模式并缓解实例特定偏差。实验表明,我们的EBM能准确模拟目标LLM,使解释器可有效识别生成特定目标输出时最具影响力的提示句子。
英文摘要
The widespread adoption of proprietary Large Language Models (LLMs) accessed strictly through closed APIs has created a critical challenge for responsible deployment: a fundamental lack of interpretability. To address this, we propose a model-agnostic, post-hoc attribution interpreter operating at the sentence level. Our approach trains an Energy-Based Model (EBM) as a surrogate to capture the LLM's internal conceptual consistency between prompts and responses. This energy landscape guides the training of a lightweight interpreter network. Uniquely, our interpreter operates as a standalone tool; once trained, it quantifies the influence of prompt sentences on a user-specified target output without requiring further API queries to the LLM. By globally training a local interpreter across diverse inputs, our framework captures broader generation patterns and mitigates instance-specific biases. Experiments demonstrate that our EBM accurately simulates the target LLM, allowing the interpreter to effectively identify the prompt sentences most influential in generating specific target outputs.
发表机构
- Sharif University of Technology(谢里夫理工大学)
机构由 AI 辅助整理,请以论文原文为准。