arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

随机采样在认知上是浅薄的:大语言模型中温度变化与模型多样性之间的维度差距

Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs

Izhar Ali

arXiv 2607.20464首次发表:更新:

发表机构

Rowan University(罗文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究语言模型重复运行答案不同时其变化能否揭示未知,比较单个模型多次运行与模型集合一次运行,通过随机矩阵测试发现单个模型跨问题结构不明显,自一致性能给出单问题不确定性,多样集合可揭示模型未知。

AI 中文摘要

当语言模型在重复运行时给出不同答案,这种变化能揭示其未知的内容吗?自一致性通过多数投票将这种变化转化为每个问题的不确定性估计。但相同的变化能揭示跨问题结构吗?我们在相同问题上比较两种情况:一个模型在τ = 1时运行100次,与24个大语言模型的集合在τ = 0时各运行一次。通过马尔琴科 - 帕斯特尔随机矩阵测试来区分信号与采样噪声。在单个模型中,五个家族和三个基准(MMLU、HellaSwag、GSM8K)中最多一个维度高于噪声。在集合中,四个特征值超过噪声边缘,而匹配难度的伯努利空值在500次蒙特卡罗抽样中最多产生一个。自一致性给出了准确的每个问题的不确定性,但没有可检测的跨问题结构;只有多样化的集合才能揭示模型未知的内容。

英文摘要

When a language model gives different answers on repeated runs, does that variation reveal what it does not know? Self-consistency turns the variation into a per-question uncertainty estimate via majority voting. But does the same variation reveal cross-question structure -- related questions flipping together, the way a diverse ensemble does? We compare two regimes on the same questions: one model run $100$ times at $τ=1$ versus an ensemble of $24$ LLMs run once each at $τ=0$. A Marchenko--Pastur random-matrix test separates signal from sampling noise on both sides. Within any single model, at most one dimension rises above noise across five families and three benchmarks (MMLU, HellaSwag, GSM8K). Across the ensemble, four eigenvalues clear the noise edge, while a matched-difficulty Bernoulli null produces at most one in $500$ Monte Carlo draws. Self-consistency gives accurate per-question uncertainty but no detectable cross-question structure; only a diverse ensemble surfaces what a model does not know.

Comments8 pages, 4 figures, 3 tables. Accepted at EIML@ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑