迈向大语言模型信念稳定性的分级度量
Toward a Graded Measure of Belief Stability in Large Language Models
浏览论文内容
中文总结 AI 辅助
本文提出分级信念稳定性度量,通过直接条件估计器利用内部表示评估LLM信念的持久性,实验显示低稳定性信念在对话挑战下行为变动更大,扩展了可靠性评估。
中文摘要 AI 辅助
大语言模型(LLMs)日益成为人们获取信息和进行推理的中介,然而事实可靠性通常是以单个判断为单位进行评估的。我们引入了分级信念稳定性,这是一种关系度量,用于衡量信念在LLM更广泛的信念系统中持续存在的程度。与单个信念概率不同,它询问的是,当某个主张与该模型的其他认知承诺一并考虑时,对该主张的支持是否持续存在。我们通过一个直接条件估计器来实现这一想法,该估计器利用内部模型表示来估计条件信念概率。在12个LLM和三个领域中,在匹配单个信念概率后,较低稳定性的信念在83.3%的模型-领域设置中表现出在对话挑战下更大的平均行为变动。因此,分级信念稳定性将可靠性评估从LLM支持某个主张的强度扩展到该信念在其更广泛的信念系统中得到支持的稳健性。
英文摘要
Large language models (LLMs) increasingly mediate how people access and reason with information, yet factual reliability is usually evaluated one judgment at a time. We introduce graded belief stability, a relational measure of how well a belief persists within an LLM's broader belief system. Unlike individual belief probability, it asks whether support for a claim persists when that claim is considered alongside the model's other epistemic commitments. We operationalize this idea with a Direct Conditional estimator that uses internal model representations to estimate conditional belief probabilities. Across 12 LLMs and three domains, lower-stability beliefs exhibit greater mean behavioral movement under conversational challenge in 83.3% of model-domain settings after matching on individual belief probability. Graded belief stability therefore extends reliability assessment beyond how strongly an LLM supports a claim to how robustly that belief is supported within its broader system of beliefs.
发表机构
- Northeastern University(东北大学)
- Santa Fe Institute(圣塔菲研究所)
机构由 AI 辅助整理,请以论文原文为准。