arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18438cs.CLcs.AIcs.LG

Relay-Bench:评估大语言模型在多领域推理链上的表现

Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains

Liam Swayne

首次发表
浏览论文内容

中文总结 AI 辅助

Relay-Bench是用于评估大语言模型在多领域推理链能力的基准测试,测试集含复合问题,涵盖多领域,模型使用无限制,鼓励利用工具,领先模型GPT-5.5(xHigh)得分43.3%,问题由两到十三个子问题组成,无需多模态输入输出。

中文摘要 AI 辅助

介绍了Relay-Bench,一个不饱和、整体的纯文本基准测试,用于衡量大语言模型在单个提示中完成各种不同领域任务的能力。领先模型GPT-5.5(xHigh)得分43.3%。测试集完全由复合问题组成,这些问题由单领域子问题串连而成,需要跨领域推理。许多问题通过提示编码和故意的上下文扩展增加了复杂性。测试领域包括视觉推理、编码、数学、信息提取、问题解决、常识和数据分析等。模型使用无限制,鼓励利用代码执行、网络搜索等工具。所有问题由两到十三个子问题组成,不需要多模态输入或输出。

英文摘要

Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete an assortment of tasks from distinct domains in a single prompt. The leading model, GPT-5.5 (xHigh), scores 43.3%. The test set entirely consists of composite problems: groups of single-domain subproblems that are strung together into challenges that require reasoning across multiple domains in combination. Many of these problems then have layers of complexity added through prompt encoding and deliberate context bloat. Domains tested include visual reasoning, coding, math, information extraction (with a focus on web search), problem-solving, general knowledge, and data analysis. No restrictions are imposed outside of the model harness, and models are explicitly encouraged to leverage code-execution, web searches, and all available tools. All problems are composed of two to thirteen subproblems and do not require multi-modal input or output.

补充信息

↑