arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Periscope:将冻结语言模型扩展到其上下文窗口之外

Periscope: Extending Frozen Language Models Beyond Their Context Window

Mohamed Eltahir, Anas Obayd, Raed Rashid, Abdulrahman Alghamdi, Abdulrahman Mousa, Abdallah Ahmed, Tanveer Hussain, Naeemullah Khan

arXiv 2610.04047首次发表:更新:

发表机构

King Abdullah University of Science and Technology (KAUST); Edge Hill University(阿卜杜拉国王科技大学; 埃奇希尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Periscope 是一种无需训练的推理方法,通过网格化文本块并利用冻结模型进行局部和跨步探测,以二次方成本扩展上下文窗口,在长文本基准上优于传统窗口读取。

AI 中文摘要

语言模型以二次方的前向传播读取长文本,在到达上下文窗口时停止,并且在达到该窗口之前,随着长度的增加而损失准确性。我们探讨在有限集合上进行决策时,是否可以将读取过程分解:哪个文档相关,哪个选项得到支持,哪个段落是证据。Periscope 是一种无需训练的推理方法,它将文本的 N 个块排列在 K×K 的网格上,其中 K=⌈√N⌉,并向冻结模型提出关于 K 个连续块的局部跨度和 K 个对整个文本进行采样的跨步跨度的相同问题,在一个 token 处读取每个答案的对数几率。每个答案取其最佳的局部和跨步得分,并通过其两个跨度对每个块进行评分,从而以零额外成本获得证据图,其峰值即为答案背后的块。对于长度为 s 个 token、块大小为 c 的文本,每次探测涉及约 √sc 个 token,因此大小为 W 的窗口可以以 s^1.5 的成本覆盖 W^2/c 个 token。该图取代了长读取。在 LongBench v2 上,仅读取地图排名最高的 K 个块(9k 个 token)即可与同一模型在 32k 到 1M token 窗口范围内的最佳窗口读取相匹配;在 InfiniteBench 上(其中位上下文为 150k 个 token),它领先最佳窗口读取 5 个百分点。同一地图在 BRIGHT 的长文档语料库上以六种方法中最佳的 NDCG@10 进行排名。每次调用仅缓存一个探测,因此一个 27B 模型可以在单个 80GB GPU 上读取 4.5M-token 的上下文,而单次传递则需要 296GB 的缓存。因此,长读取只需要一个能容纳模型的 GPU,而不是一个能容纳文本的 GPU。

英文摘要

A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which passage is the evidence. Periscope, a training-free inference method, arranges the $N$ chunks of a text on a $K{\times}K$ grid with $K{=}\lceil\sqrt{N}\rceil$ and asks a frozen model the same question about $K$ local spans of consecutive chunks and $K$ strided spans that sample the whole text, reading the log-odds of every answer at one token. Each answer takes its best local and strided score, and scoring every chunk by its two spans gives an evidence map at no further cost, whose peak is the chunk behind the answer. Every probe is about $\sqrt{sc}$ tokens for a text of $s$ tokens and chunk size $c$, so a window of $W$ tokens reaches $W^{2}/c$ tokens at $s^{1.5}$ cost. The map replaces the long read. On LongBench v2, reading only the $K$ chunks the map ranks highest, 9k tokens, matches the same model's best window read across windows from 32k to 1M tokens, and on InfiniteBench, where the median context is 150k tokens, it leads the best window read by 5 points. The same map ranks BRIGHT's long-document corpora with the best NDCG@10 of six methods. Each call caches only one probe, so a 27B model reads 4.5M-token contexts on one 80GB GPU, where a single pass would need 296GB of cache. A long read then needs a GPU that holds the model, not one that holds the text.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑