AI 中文总结
研究指出生成式人工智能使科学中凭说服力充当证据的老问题重现,将想象与实际结果差距定义为虚幻证据,发现分辨率等不增证据,单个结果证据有上限,已发表真实发现比例会回落,解决办法是拓宽系统范围并按培根方法测量。
AI 中文摘要
四个世纪前,弗朗西斯·培根警告人们要警惕对自然的臆测,即仅凭少量事实就匆忙归纳,并提出了“否定事例表”:检查某种特性在不应出现的地方是否未出现。要求是看似有说服力本身不应算作证据。从那时起,科学一直宣称有此要求,但实际上却让说服力充当证据。生成式人工智能消除了这一困难,一个古老的错误以前所未有的规模和速度卷土重来。我们发现问题不在于证据变弱,而在于如何计算惊喜程度。观察者惊叹于一个有说服力的输出是众多可能性中的一个单点命中,但系统实际能达到的只是该范围的一小部分。我们将想象的广度与实际达到的狭窄度之间的差距称为虚幻证据,并将其形式化为一个量,该量还吸收了研究过程中的试错和数据泄露。由此得出三点:更高的分辨率和更高的流畅度不会增加证据;单个结果所能承载的证据有上限,无论是打磨输出还是让生成系统自评都无法超越;已发表发现中真实的比例会回到观察前的水平。解决方法在于真正拓宽系统能达到的范围,并测量在目标不存在时是否仍会出现有说服力的输出——用概率语言重述的培根“否定事例表”。在一个说服力变得廉价的世界里,科学的可信度不在于更有说服力的输出,而在于表明它们不可能偶然出现的程序。
英文摘要
Four centuries ago Francis Bacon warned against the anticipations of nature, hasty generalization that wins assent on a few facts, and set against it the table of absence: checking that a property fails to appear where it should not. The demand was that looking convincing should not, on its own, count as evidence. Science has professed that demand ever since, while in practice letting persuasiveness do the work of evidence. It could be let to do so because making something persuasive was itself hard. Generative AI removes that difficulty, and an old error returns on a scale and at a speed it never had before. We locate the problem not in evidence growing weaker but in how surprise is counted. An observer marvels at a convincing output as a single point hit among a vast range of possibilities, yet what a system can actually reach is a small part of that range. The gap between the breadth imagined and the narrowness actually reached is what we call phantom evidence, and we formalize it as one quantity that also absorbs the trial and error and the data leakage a research process adds. Three things follow. Higher resolution and greater fluency add no evidence. The evidence a single result can carry has a ceiling that neither polishing the output nor letting a generative system grade itself can exceed. And the fraction of published findings that are true falls back to what it was before anything was observed. The prescription lies in the same place: genuinely widen what a system can reach, and measure whether convincing outputs still appear when the target is absent -- Bacon's table of absence, restated in the language of probability. In a world where the persuasive has become cheap, the credibility of science rests not on more convincing outputs but on procedures that show they could not have arisen by chance.