arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2607.26448cs.CLcs.AI

面向长上下文语言模型的可合并模型侧聚合状态

Mergeable Model-Side Aggregation States for Long-Context Language Models

Dachuan Song, Junyu Yin, Zechen Hu, Xuan Wang

AI总结:

针对长上下文语言模型在非加性集合聚合任务中的性能缺陷,提出带HLL草图状态的模型侧聚合接口,该接口可合并跨段状态,在多任务上较基线方法实现显著性能提升。

AI中文摘要:

长上下文语言模型的一个已知局限是,随着上下文长度增长,其在非加性、基于集合的聚合任务中的性能愈发不可靠。这类任务广泛存在于日志、程序输出、表格及多轮对话中,例如基数估计、集合关系判断、分组统计等。为提供这些任务所需的聚合状态,我们提出一种模型侧聚合接口,该接口在冻结语言模型的同时维护紧凑的基于哈希的HyperLogLog(HLL)草图状态。模型处理上下文时,提取器会将每个相关记录映射为规范标识,该标识经哈希后更新HLL状态。这些状态可跨上下文段合并,或直接读出用于下游推理,避免了额外的生成-执行-返回循环。我们通过将HLL状态大小设为2 KiB(2048个寄存器)来验证所提方法,该大小不随上下文长度或集合基数增长。在涉及100万条记录的基数估计实验中,平均相对误差为1.6%;在独立合并测试中,由多达256个段构建的状态,其读出结果与对同一数据流的单次遍历完全一致。在来自174个源窗口的3969个“聚合后推理”任务上,固定预算接口在Gemma 4(31B,BF16)上达到99.2%的准确率,而精确聚合下准确率为100.0%,配对差距为0.8个百分点(95%窗口聚类置信区间:0.5-1.3个百分点)。在匹配的174个条目上,我们的方法相比直接全上下文推理,在Qwen上提升63.2个百分点,在Gemma上提升56.3个百分点;相比思维链(CoT)推理,对应提升分别为60.9和63.2个百分点。在固定的1200任务Oolong-Synth子集上,我们的方法在Qwen上达到91.1%,在Gemma上达到99.3%。代码可在指定URL获取。

英文摘要:

A known limitation of long-context language models is their increasingly unreliable performance in non-additive, set-based aggregation as context length grows. Examples include cardinality estimation, set relationships, and grouped statistics, which widely exist in logs, program outputs, tables, and multi-turn conversations. To provide the aggregation state required by these tasks, we introduce a model-side aggregation interface that maintains compact Hash-based HyperLogLog (HLL) sketch states alongside a frozen language model. While the model processes the context, an extractor maps each relevant record to a canonical identity. The identity is then hashed and updates the HLL state. These states can be merged across context segments and/or read out directly for downstream reasoning, avoiding an additional generate-execute-return cycle. We validate the proposed approach by setting the HLL state size as 2 KiB (2,048 registers), which does not increase with context length or set cardinality. In a distinct-count experiment involving one million records, the mean relative error was 1.6%. In a separate merge test, states built from as many as 256 segments produced exactly the same readout as a single pass over the same stream. On 3,969 aggregate-then-reason tasks from 174 source windows, the fixed-budget interface reached 99.2% accuracy on Gemma 4 (31B, BF16), compared with 100.0% under exact aggregation; the paired gap was 0.8 percentage points (95% window-cluster CI: 0.5-1.3 points). On a matched set of 174 items, our method improved over direct full-context reasoning by 63.2 points on Qwen and 56.3 points on Gemma. The corresponding gains over chain-of-thought (CoT) reasoning were 60.9 and 63.2 points, respectively. On a fixed 1,200-task Oolong-Synth subset, our method reached 91.1% on Qwen and 99.3% on Gemma. Code is available at https://github.com/songdc98/sketchops.

↑