发表机构
University of Toronto; Vector Institute(多伦多大学; 向量研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ArbiGraph作为基准生成器,将任务表示为带Python求解器的自然语言问题,通过类型化中间状态组合任务。用多种任务类别实例化并评估智能体,发现其在孤立任务和复杂相关任务上表现不同,揭示了单任务评估无法发现的问题。
AI 中文摘要
我们引入了ArbiGraph,这是一种基准生成器,用于评估工具辅助语言智能体能否在扩展推理工作流程中保留、更新、组合和丢弃与任务相关的上下文。ArbiGraph将每个任务表示为带有可执行Python求解器的自然语言问题,并通过类型化中间状态组合任务,在此实例化为标量和列表值。这种设计实现了可控任务图,其长度、依赖结构、干扰项数量和值类型可以变化,同时保持精确的自动验证。我们用数学、GSM风格的文字问题和Python跟踪任务类别实例化ArbiGraph,并在四种拓扑结构上评估Qwen3.5 - 27B工具辅助智能体。结果表明,在孤立任务上准确率高,但在更复杂的相关任务上大幅下降:在相关数学任务的分支链上准确率下降高达33.3%。这表明ArbiGraph揭示了仅通过单任务评估无法看到的失败情况。我们的代码、生成的数据集和评估结果可在该https URL获取。
英文摘要
We introduce ARBIGRAPH, a benchmark generator for evaluating whether tool-assisted language agents can retain, update, compose, and discard task-relevant context across extended reasoning workflows. ARBIGRAPH represents each task as a natural-language problem with an executable Python solver, and composes tasks through typed intermediate states, instantiated here as scalar and list values. This design enables controllable task graphs whose length, dependency structure, distractor count, and value type can be varied while preserving exact automatic verification. We instantiate ARBIGRAPH with math, GSM-style word-problems, and Python-tracing task categories, and evaluate a Qwen3.5-27B tool-assisted agent across four topologies. The results show high accuracy on isolated tasks but substantial degradation on more complex dependent tasks: accuracy drops by up to 33.3% on branching chains of dependent math tasks. This shows that ARBIGRAPH exposes failures that are not visible from single-task evaluation alone. Our code, generated datasets, and evaluation results are available at https://github.com/pavelgolikov/ArbiGraph.git