arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37499cs.CL

谁让档案变暖了?LLMs 高估历史温暖程度

Who Warmed the Archives? LLMs Overestimate Historical Warmth

Claudiu Creanga, Liviu P. Dinu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究评估 LLM 从历史档案提取气候指数时的跨世纪偏差,发现六个 LLM 均存在随年份增长的温暖偏差,仅凭相关性不足以验证其可靠性。

中文摘要 AI 辅助

历史档案是延长仪器气候记录回溯时间的未充分利用的资源,而大语言模型(LLMs)提供了一种提取气候学家手工推导的指数的方法。除了衡量系统提取该信号的效果外,我们还检查了其误差是否可用于跨世纪比较,因为良好的相关性分数并不能排除系统性的、与时代相关的偏差。通过比较词汇基线、微调的历史变换器以及 LLM 提示方法在五个世纪的德语文本上对 Pfister 温度指数的表现,词汇方法击败了我们测试的所有微调变换器,包括一个从头开始在历史德语上预训练的模型(r=-0.016)。我们测试的所有六个 LLM(Gemini 2.5 Flash、GPT-5-mini、DeepSeek v4 Flash、Claude Sonnet 4.6、Qwen3.7-Plus、Kimi-K2.6-Fast)都表现出随日历年份增长而增强的温暖偏差,且每个模型中的符号相同(斜率+0.13至+0.34/世纪,p<0.01)。该效应在规模上适中(r平方约等于0.01至0.05),但在六个独立开发的模型中保持一致。六个模型中相关性最佳的 Gemini 2.5 Flash 在误差加倍的情况下匹配了最佳词汇相关性(r=0.32)。从引文中剥离显式日期和日历时代标记的消融实验使这一趋势基本保持不变,这支持了模型持有不合时宜的当下先验,而非正确推断引文时代。因此,仅凭相关性不足以将 LLM 验证为历史气候指数预言机。

英文摘要

Historical archives are an under-used source for extending the instrumental climate record backward in time, and LLMs offer a way to extract the indices climatologists derive by hand. Beyond measuring how well systems extract this signal, we check whether their errors are safe to use for cross-century comparison, since a good correlation score does not rule out systematic, era-linked bias. Comparing lexical baselines, fine-tuned historical transformers, and LLM prompting on the Pfister temperature index across five centuries of German text, lexical methods beat every fine-tuned transformer we test, including one pretrained from scratch on historical German (r=-0.016). All six LLMs we test (Gemini 2.5 Flash, GPT-5-mini, DeepSeek v4 Flash, Claude Sonnet 4.6, Qwen3.7-Plus, Kimi-K2.6-Fast) show a warm bias that grows with calendar year, with the same sign in every model (slopes +0.13 to +0.34/century, p<0.01). The effect is modest in size (r-squared approx equal to 0.01 to 0.05) but consistent across six independently developed models. The best-correlated of the six, Gemini 2.5 Flash, matches the best lexical correlation (r=0.32) at double the error. An ablation stripping explicit dates and calendar-era markers from the quotes leaves this trend essentially unchanged, favoring an anachronistic present-day prior over the model correctly inferring the quote's era. Correlation alone is thus insufficient for vetting an LLM as a historical-climate-index oracle.

发表机构

  • University of Bucharest(布加勒斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

↑