arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23178cs.CL

Chronologic:衡量语言模型表征过去的能力

Chronologic: Measuring Language Models' Ability to Represent the Past

Ted Underwood, Ziliang Qiu, Sarah Griebel, Laura K. Nelson, Edwin Roland, Wenyi Shang, Matthew Wilkens

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出Chronologic基准,通过历史文本成对比较评估语言模型对1831-1930年英语语境的表征能力,发现生成任务更难,且当前模型表现尚不完美但有进展。

中文摘要 AI 辅助

语言模型是研究过去的吸引人的工具。但为了信任模型提供的证据,研究人员需要知道其回答是否符合所代表的时期。验证具有挑战性,因为这不是普通人通常执行的任务,而且许多问题有多个正确答案。我们利用历史文本开发了一个基准,用于衡量模型对1831-1930年间英语语境的表现,依靠成对比较多个真实答案和强干扰项,以适当分级的方式对最难的问题进行评分。我们发现生成任务比判别任务更难;事实上,推理模型通常能辨别自己生成答案的弱点。虽然仅基于历史文本预训练的模型在按答案可能性评估时领先,但在自由生成方面无法与商业模型竞争。我们测试的所有模型都尚未完全令人信服地表现历史语境,但朝着这一目标的进展是明显的。

英文摘要

Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct answers. We use historical texts to develop a benchmark for a model's representation of English-language contexts 1831-1930, relying on pairwise comparisons to multiple ground truths and strong distractors to score the hardest questions in an appropriately graduated way. We find that generative tasks are harder than discriminative ones; in fact, reasoning models can typically discern the weakness of their own generated answers. While models pretrained exclusively on historical text lead the pack when evaluated by answer likelihood, they cannot compete with commercial models in free generation. None of the models we tested represent historical contexts in a fully persuasive way yet, but progress toward that goal is evident.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • University of British Columbia(不列颠哥伦比亚大学)
  • University of Missouri(密苏里大学)
  • Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑