arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型中的信念与行为

Beliefs and Behavior in Language Models

Alex Smolin, Bryan Wilder

arXiv 2609.07943首次发表:更新:

发表机构

Toulouse School of Economics; Carnegie Mellon University(图卢兹经济学院; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出实证方法,通过推断语言模型输出中的潜在信念变量来预测其行为,发现高能力模型可被有效描述为持有信念,并探讨了信念的测量、决策规则遵循及推理中的演变。

AI 中文摘要

关于诸如信念或欲望之类的抽象概念是否能有效描述大型语言模型(LLMs)的行为,存在显著的不确定性。除了这一问题本身固有的科学意义外,这些潜在量常被用来向用户解释LLMs的行为,或用于定义和评估与意图相关的有害行为。然而,我们目前缺乏一种系统性的方法来检验“信念”等概念是否适用于LLMs,以及它们是否可能成为将模型与人类利益对齐的有益要素。我们提出了一种实证研究此类问题的方法,即询问从LLMs输出中推断出的单一潜在变量——被解释为信念程度——是否能使观察者对LLMs如何响应新提示做出可解释的预测。我们发现,高能力模型可以被有效地描述为持有信念,并且总体而言,基于推断出的潜在信念对模型输出的可预测性,与模型能力的总体趋势相一致。基于这些发现,我们提供了实证策略,以研究如何测量LLMs中的信念、LLMs在多大程度上遵循指定的决策规则或收益,以及信念在单个LLM实例的推理过程中如何演变。

英文摘要

There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In addition to the inherent scientific interest of this question, these latent quantities are often invoked to explain the behavior of LLMs to users or to define and evaluate harmful behaviors which are relative to intent. Nevertheless, we currently lack a means to systematically test whether concepts like "belief" are well-applied to LLMs, and hence whether they are likely to be fruitful ingredients of attempts to align models with human interests. We propose an approach for empirically studying such questions, asking whether a single latent variable inferred from the LLMs' outputs -- interpreted as a degree of belief -- allows an observer to make interpretable predictions of how the LLMs' will respond to new prompts. We find that highly capable models are usefully described as holding beliefs and that, generally, the predictability of model outputs based on an inferred latent belief tracks overall trends in model capability. Building on these findings, we provide empirical strategies to study how beliefs in LLMs can be measured, the extent to which LLMs comply with instructed decision rules or payoffs, and how beliefs evolve within individual instances of an LLM over the course of reasoning.

Comments33 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑