arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13529cs.LGcs.AIcs.CLcs.SC

可扩展神经符号模型下的生成式可解释性

Generative Interpretability via Scalable Neuro-Symbolic Models

Xiaocong Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对大模型在智能体系统中输出不可逆行动的安全问题,提出生成式可解释性架构,使推理暴露可干预的语义检查点,并以神经符号模型实现。

中文摘要 AI 辅助

随着大语言模型的应用从聊天机器人转向智能体系统,其输出成为对现实具有不可逆后果的行动,现有的人工智能可解释性研究范式(事后可解释性)在结构上不足以支持安全可信的模型部署:它只能在事后解释行为,却无法在推理计算提交输出之前对其进行审计或干预。因此,我们主张转向“生成式可解释性”,这是一种架构属性,在该属性下,模型的推理过程天然地暴露语义上有意义的检查点,这些检查点是人类可理解的且适合因果干预。我们展示了生成式可解释性相较于其他可解释性研究范式的优势,并提出了神经符号模型作为其具体实例化。

英文摘要

As the use of Large Language Models moves from chatbots into agentic systems, where outputs become actions with irreversible consequences on reality, the existing paradigm on AI Interpretability research, post-hoc interpretability, is structurally inadequate for safe and trustworthy model deployment: it explains behavior after the fact but cannot audit or intervene in an inference computation before it commits to an output. We therefore argue for a shift toward \emph{generative interpretability}, an architectural property under which a model's inference pass natively exposes semantically meaningful checkpoints that are human-understandable and amenable to causal intervention. We show the merits of generative interpretability as comparison to other interpretability research paradigms, and propose Neuro-Symbolic Models as a concrete instantiation.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑