发表机构
University of Colorado, Boulder(科罗拉多大学博尔德分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探究LLMs是否具备格赖斯式退却的要素,发现其激活可编码知识边界与指称特异性信号,但生成时未协调,为格赖斯式对齐提供了基础。
AI 中文摘要
当被问及知识边界之外的实体时,大型语言模型(LLMs)通常会编造看似合理的细节,而非退回到更安全、更具普遍性的表述。我们通过格赖斯视角来阐释这一失败:对指称不确定的合作说话者会在特异性层级上后退,以牺牲信息量为代价换取真实性。我们探究LLMs是否具备执行这种退却的要素。使用基于T-REx的基准(该基准会改变实体熟悉度和指称特异性),我们探究模型以回答两个问题:(i)它们的激活是否编码了指称是否落在知识边界内;(ii)它们是否预期即将生成的指称的特异性?我们发现两个问题的答案均为肯定,但两种信号在生成过程中未得到协调。即使实体对模型而言是未知的,且即便提供了正确的通用替代方案,模型也压倒性地偏好具体的指称。格赖斯式退却的基础已然存在,但据此采取行动的策略却不存在。我们将这些发现视为迈向格赖斯式对齐的第一步,即训练或引导目标在生成过程中将知识边界意识与指称特异性相结合。
英文摘要
When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to safer, more general claims. We frame this failure through a Gricean lens: a cooperative speaker who is uncertain about a referent retreats up the specificity hierarchy, trading informativeness for truthfulness. We ask whether LLMs have the ingredients to perform this retreat. Using a T-REx-based benchmark that varies entity familiarity and referent specificity, we probe models to answer two questions: (i) do their activations encode whether a referent falls inside the knowledge boundary, and (ii) do they anticipate the specificity of the referent they are about to generate? We find that the answer to both is yes, but the two signals are not reconciled in generation. Models overwhelmingly prefer specific referents even when the entity is unknown to them, and do so even when offered correct generic alternatives. The substrate for a Gricean retreat is present, but the policy that would act on it is not. We position our findings as a first step toward Gricean alignment, training or steering objectives that couple knowledge-boundary awareness to referent-specificity during generation.