arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

定价引导:语言模型能否生成未来研究想法?

Priced Guidance: Can Language Models Generate Future Research Ideas?

Kaiyue Wen, Tengyu Ma, Percy Liang

arXiv 2610.04976首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出定价引导框架,通过压缩成本度量语言模型生成未来研究想法的能力,并在87篇论文上验证,发现Fable 5.1最优,集成可进一步提升性能。

AI 中文摘要

我们通过压缩的视角评估语言模型生成新颖研究想法的能力。我们的目标是给出一个下界,即语言模型在没有任何提示的情况下生成未来研究想法本质的潜在极小概率。我们不是通过昂贵的重复采样来估计这一概率,而是采用我们的定价引导框架来衡量压缩成本:即需要多少额外的信息比特来引导模型恢复目标想法。我们证明,如果模型在期望上最多使用K比特的引导就能恢复目标想法,那么它在没有任何引导的情况下生成该想法的概率至少为$2^{-K}$。在我们的框架中,语言模型(称为生成器)可以提出一系列多项选择题,并指定可能答案上的概率分布。一个能够访问目标想法的引导者(也是一个语言模型)选择答案。如果所选答案的先验概率为p,则生成器支付$-\u005c\u005clog_2 p$比特。生成器的目标是以最小的累积成本生成与目标想法本质匹配的想法。该累积成本,加上一个加性常数,等于引导者发送的信息比特数。使用这种方法,我们评估了五个生成器大语言模型(Opus 5、Fable 5.1、GPT-6 Astra、GPT-5.6 Sol 和 GLM 5.3)在87篇近期高质量深度学习论文中的核心想法上,并使用大语言模型作为评判者来判断生成的想法是否在核心研究目标和定义机制方面与目标匹配。Fable 5.1实现了最低的中位压缩成本69.9比特,远低于gzip对目标想法摘要进行无损压缩的中位5712比特。Fable 5.1、Opus 5和Astra的均匀集成进一步将中位压缩成本降至55.8比特,并将生成概率下界提高了18000倍。

英文摘要

We evaluate language models' capability to generate novel research ideas through the lens of compression. We aim to lower-bound the potentially tiny probability that a language model generates the essence of a future research idea without any hints. Rather than estimate this probability through expensive repeated sampling, our Priced Guidance framework measures the compression cost: how many additional bits of information are needed to guide the model to recover the target idea. We prove that if the model can recover the target idea with at most K bits of guidance in expectation, then it can generate the idea without any guidance with probability at least $2^{-K}$. In our framework, the language model, called the generator, can pose a sequence of multiple-choice questions and specify a probability distribution over possible answers. A guide, which is a language model with access to the target idea, selects answers. If a selected answer has prior probability p, the generator pays $-\log_2 p$ bits. The generator aims to produce an idea that matches the essence of the target idea with minimal cumulative cost. This cumulative cost equals, up to an additive constant, the number of bits of information sent by the guide. Using this methodology, we evaluate five generator LLMs (Opus 5, Fable 5.1, GPT-6 Astra, GPT-5.6 Sol, and GLM 5.3) on the core ideas in 87 recent high-quality deep learning papers and we use LLM as judge to determine whether the generated idea matches the target in terms of the central research object and defining mechanism. Fable 5.1 achieves the lowest median compression cost at 69.9 bits, substantially lower than gzip's median of 5,712 bits for losslessly compressing the summary of the target idea. A uniform ensemble of Fable 5.1, Opus 5, and Astra further reduces the median compression cost to 55.8 bits and improves the generation probability lower bound by 18,000 times.

Comments80 pages, 12 figures. Code available at https://github.com/WhenWen/priced-guidance

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑