arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38137cs.CL

LongHarness Bench:对长上下文推理的语言模型框架进行压力测试

LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

Quang Hieu Pham, Thuy Duong Nguyen, Jocelyn Qiaochu Chen, Xi Ye

首次发表
浏览论文内容

中文总结 AI 辅助

提出LongHarness Bench基准,用于评估长上下文语言模型框架的有效性与效率,发现相同模型在不同框架下效率差异显著,最佳组合宏平均准确率仅68%。

中文摘要 AI 辅助

语言模型(LM)框架通过额外的计算使语言模型能够在长上下文中高效运行。然而,现有的长上下文评估不足以区分现代框架,这体现在各框架间准确率趋于饱和以及评估成本大致相似。在本文中,我们引入了一个基准,用于评估长上下文框架的有效性和效率。我们的任务需要多样化的检索策略,包括词汇搜索和语义匹配,以及对全局和局部上下文的策略性和适应性推理。大部分上下文在语义上相关,但每一步只有一小部分是有用的,这既构成了一个具有挑战性的搜索问题,也在不同处理策略之间产生了不同的准确率-成本权衡。例如,一项任务要求利用分散在文档中的证据识别满足多个条件的每个人;策略性地先检查最具选择性的条件,可以在验证其余条件之前缩小搜索范围。我们使用四种最先进的框架评估了多个前沿语言模型家族。我们的基准即使对强大的模型-框架组合也仍然具有挑战性:最佳组合在四个评估套件上达到了68%的宏平均准确率。更重要的是,我们发现相同的底层模型在不同的框架下可以表现出显著不同的效率。我们的结果确立了效率作为长上下文评估的一个重要维度,并为开发能够策略性地而非穷尽性地处理上下文的框架提供了测试平台。

英文摘要

Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated accuracy across harnesses and largely similar evaluation costs. In this paper, we introduce a benchmark for evaluating both the effectiveness and efficiency of long-context harnesses. Our tasks require diverse retrieval strategies, including lexical search and semantic matching, together with strategic and adaptive reasoning over global and local context. Much of the context is semantically relevant but only a small subset is useful at each step, creating both a challenging search problem and different accuracy--cost tradeoffs across processing strategies. For example, one task requires identifying every person satisfying several conditions using evidence scattered across documents; strategically checking the most selective condition first can narrow the search before verifying the remaining conditions. We evaluate multiple families of frontier language models with four state-of-the-art harnesses. Our benchmarks remain challenging even for strong model--harness combinations: the best reaches 68\% macro-average accuracy across four evaluation suites. More importantly, we find that the same underlying model can exhibit markedly different efficiency under different harnesses. Our results establish efficiency as an important axis for long-context evaluation and provide a testbed for developing harnesses that process context strategically rather than exhaustively.

发表机构

  • University of Alberta(阿尔伯塔大学)
  • Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑