arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02549cs.CL

评估大型语言模型在时间抽取任务中的多维泛化能力

Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks

Fahmid Shahriar Iqbal, Ritam Dutt, Soumitra Das, Arnav Verma, Sagnik Ray Choudhury

首次发表
浏览论文内容

中文总结 AI 辅助

本研究系统评估了多种大型语言模型在时间与事件表达抽取任务中跨四个维度的泛化能力,发现基础任务性能通常预示更好泛化但在显著分布变化下减弱,归纳提示表现最一致,而规模、架构及演绎和溯因提示策略的收益不均,表明泛化无法从单一维度可靠预测。

中文摘要 AI 辅助

时间和事件表达抽取是基础的时间推理任务,但由于标注歧义、领域敏感性和模型行为不稳定,该问题仍然困难。现有评估侧重于域内性能,对分布变化下的可靠性提供的洞察有限。我们评估了跨家族、架构和推理策略的多种模型配置在四个泛化维度上的表现,考察了从基础性能的迁移、跨维度相关性以及规模、架构和提示的影响。这提供了对提示型LLM在时间和事件表达抽取任务中如何泛化的系统性研究。我们发现,强的基础任务性能通常预示着更好的泛化能力。然而,这种关系在显著分布变化下会减弱。归纳提示在领域迁移、对抗扰动、组合性和长度增加方面表现最一致,而规模、架构以及演绎和溯因提示策略的收益则不均匀且具有维度特异性。我们得出结论,LLM在时间抽取任务中的泛化不能仅从任何单一维度预测,也不能可靠地从域内或单维度评估中推断,这凸显了需要跨维度泛化的推理策略。

英文摘要

Time and event expression extraction are fundamental temporal reasoning tasks, but the problem remains difficult due to annotation ambiguity, domain sensitivity, and unstable model behavior. Existing evaluations focus on in-domain performance, offering limited insight into reliability under distribution shifts. We evaluate multiple model configurations across families, architectures, and reasoning strategies over four dimensions of generalization, examining transfer from base performance, cross-dimensional correlations, and the effects of scale, architecture, and prompting. This provides a systematic study of how prompted LLMs generalize in time and event expression extraction tasks. We find that strong base-task performance generally predicts better generalization. However, this relationship weakens under substantial distribution shifts. Inductive prompting performs most consistently across domain shift, adversarial perturbations, compositionality, and length increase, while gains from scale, architecture, and deductive and abductive prompting strategies are uneven and dimension-specific. We conclude that LLM generalization in temporal extraction tasks cannot be predicted from any single dimension alone and cannot be reliably inferred from in-domain or single-dimension evaluations, highlighting the need for reasoning strategies that generalize across dimensions.

发表机构

  • University of North Texas(北德克萨斯大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑