arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29245cs.CLcs.AI

没有免费的午餐:随着语料库增长,语料库任务复杂度至关重要

No More Free Lunch: Corpus Task Complexity Matters as Corpora Grow

  • UC Berkeley(加州大学伯克利分校)
  • Allen Institute for AI(艾伦人工智能研究所)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Prasann Singhal, Amanda Bertsch, Jacob Steinhardt, Sewon Min

AI总结:

本文提出语料库任务复杂度(CTC)概念,区分难度随语料库线性或二次增长的任务,引入10个高CTC任务,发现其挑战长上下文模型并逆转低CTC结论,发布CTC-Bench套件。

AI中文摘要:

给定一个大型语料库,人们可能提出的问题各不相同——从“第一例人类心脏移植是什么时候?”到“这篇文献中所有相互矛盾的声明是什么?”——但是什么使得某些问题比其他问题更具挑战性?在这项工作中,我们定义了语料库任务复杂度(CTC)的概念,该概念通过任务难度随语料库规模增长的方式来表征任务;例如,检索查询只需要对语料库进行一次线性扫描,而发现矛盾则需要检查数量呈二次方增长的声明对。观察到先前的工作主要只研究了难度随语料库规模线性增长的任务,我们称之为低CTC任务,我们引入了10个属于高CTC类别的新任务,其难度随语料库规模呈二次方或更高次增长。我们发现,对于长上下文语言模型(LCLMs),高CTC任务不仅在较长上下文中平均变得更加具有挑战性,而且它们逆转了许多仅从低CTC评估中得出的建模结论。例如,高效的块稀疏和混合注意力方法在低CTC任务上始终与全注意力性能相匹配,但在高CTC任务上则退化得更严重。因此,大规模语料库的高CTC推理仍然是一个开放的挑战,因为全注意力在扩展时成本过高,这激励了未来在这些任务上的研究。我们发布了我们的代码、数据和包含22个任务的套件(CTC-Bench),以促进该领域的未来研究。

英文摘要:

Given a large corpus, the questions one might ask can vary -- from "When was the first human heart transplant?" to "What are all the contradictory claims in this literature?" -- but what makes some questions more challenging than others? In this work, we define a notion of Corpus Task Complexity (CTC) that characterizes tasks by how their difficulty grows with corpus size; for instance, a retrieval query only requires a single linear pass over a corpus, while finding contradictions requires checking a quadratically growing set of claim pairs. Observing that prior work has largely only studied tasks whose difficulty grows linearly with corpus size, which we call low CTC tasks, we introduce 10 new tasks belonging to a class of high CTC whose difficulty grows quadratically or more in corpus size. We find that high-CTC tasks not only grow much more challenging on average at longer contexts for LCLMs, they reverse many modeling conclusions drawn solely from low-CTC evaluations. For instance, efficient block-sparse and hybrid attention approaches consistently match full attention performance on low-CTC tasks, but degrade much more on high-CTC tasks. Large-corpus high-CTC reasoning thus remains an open challenge as full attention is too costly to scale, motivating future research on these tasks. We release our code, data, and 22-task suite (CTC-Bench), to facilitate future research in this area.

补充信息

↑