arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

划分、提示、聚合:语言模型中的统计自一致性

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

Patrik Wolf, Thomas Kleine Buening, Andreas Krause, Celestine Mendler-Dünner

arXiv 2607.15277首次发表:更新:

AI 中文总结

研究语言模型估计是否遵循自一致性原则,用二叉树划分总体,以子总体描述提示模型并聚合结果,发现广泛违反一致性,揭示宏观谬误模式,表明模型虽有子总体知识但未有效用于总体估计,建立了评估语言模型的新准则。

AI 中文摘要

上下文学习通常被解释为一种条件推理形式,其中提示指定上下文,模型输出被视为相应条件分布的估计。若此解释成立,语言模型估计应满足基本概率恒等式。本文研究语言模型估计在多大程度上遵循这种自一致性原则。使用二叉树递归划分总体为更细粒度子总体,用子总体描述提示语言模型,聚合结果估计并跨不同粒度划分比较。研究发现广泛违反基本一致性属性,深入研究角色提示揭示一种宏观谬误模式,该效应在树结构和估计任务变化中持续,可通过隐式提示部分恢复。这些发现表明模型拥有相关子总体知识但未可靠传播到总体估计中,为评估语言模型建立了统计自一致性这一无饱和、无参考的标准。

英文摘要

In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a context and the model's output is treated as an estimate of the corresponding conditional distribution. If this interpretation holds, then LLM estimates should satisfy basic probabilistic identities. In particular, the law of total probability asserts that prior-weighted conditional distributions aggregate into population-level marginals over any valid partition of the population. In this work, we investigate to what extent LLM estimates adhere to this self-consistency principle. We use binary trees as an evaluation scaffold to recursively partition a population into increasingly fine-grained subpopulations. We then prompt LLMs with verbalized subpopulation descriptions in context, aggregate the resulting estimates back into population-level estimates, and compare them across partitions of varying granularity. Applying this protocol across problem domains and state-of-the-art frontier models, we show widespread violations of basic consistency properties. An in-depth study of persona prompting reveals a pattern we call the macro fallacy: estimates reconstructed from more fine-grained subpopulation responses are often better aligned with human reference data than direct population-level estimates. This effect persists across variations in tree structure and estimation task, and can be partially recovered through implicit prompting. Together, these findings suggest that models possess relevant subpopulation knowledge but do not reliably propagate it into aggregate estimates. This gap establishes statistical self-consistency as an unsaturated, reference-free criterion for evaluating LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑