arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29251cs.AIcs.CL

策略即代码:用于CAR-bench快速推理可靠性的协程桥接框架

Policy as Code: A Coroutine-Bridge Harness for Fast-Reasoning Reliability on CAR-bench

Ivan Matveev

首次发表
浏览论文内容

中文总结 AI 辅助

提出协程桥接框架,让智能体仅生成可阻塞恢复的Python程序,将模型调用与工具往返解耦,在CAR-bench上以60% Pass^3赢得赛道2,延迟低至3.14秒,缓存命中率达78%。

中文摘要 AI 辅助

CAR-bench评估了工具使用智能体在现实世界不确定性下是否保持可靠,它在评估器内部执行每个工具,使得每次工具结果交换都是智能体的一次独立往返。传统的下一步行动智能体可以批量并行调用工具,但一串依赖调用每次结果往返都需要一次模型调用。我们提出了一种协程桥接框架,其中模型的唯一动作是发出一个Python程序,该程序在评估器工具交换过程中原地阻塞和恢复。这将模型调用与工具往返解耦:在公开测试集上,智能体每个任务中位数仅需两次模型调用,而传统方式需要七次智能体轮次,在Cerebras gpt-oss-120b上以中位数1.8秒的模型延迟解决完整的多轮任务。由于动作表面是可执行代码,确定性的CAR-bench策略直接编码为工具层中的逻辑,而非提示规则,以零推理成本强制执行合规性。在官方隐藏评估中,该框架以60.0%的Pass^3赢得了赛道2,是组织者基线的4.5倍,且在所有高于该基线的参赛作品中具有最低的估计成本和最快的任务中位延迟(3.14秒);相同的未更改框架在Open赛道中于GPT-5.5上复现了相同的60.0% Pass^3,与前沿模型智能体相当。一个静态提示,在尾部附加每个任务的状态,在调用和任务之间保持字节一致:冻结的提交提示从缓存中提供了78%的输入令牌(其热尾部为86.6%),而在三周开发语料库中,提示编辑反复重置缓存,缓存命中率为73%。这使少调用设计进一步将名义输入计算压缩到很小一部分。

英文摘要

CAR-bench evaluates whether tool-using agents stay reliable under real-world uncertainty, executing every tool inside the evaluator so that each tool-result exchange is a separate agent round-trip. A conventional next-action agent can batch parallel tool calls, but a chain of dependent calls costs it one model call per round of results. We present a coroutine-bridge harness in which the model's only action is to emit a Python program that blocks and resumes in place across evaluator tool exchanges. This decouples model invocation from tool round-trips: on the public test split the agent uses a median of two model calls against seven agent turns per task, resolving a full multi-turn task in a median of 1.8 s of model latency on Cerebras gpt-oss-120b. Because the action surface is executable code, deterministic CAR-bench policies are encoded directly as logic in the tool layer rather than as prompt rules, enforcing compliance at zero reasoning cost. On the official hidden evaluation the harness won Track 2 with 60.0% Pass^3, 4.5x the organizer baseline, at the lowest estimated cost and the fastest median task latency (3.14 s) of any entry scoring above that baseline; the same unchanged harness reproduced an identical 60.0% Pass^3 on GPT-5.5 in the Open track, matching frontier-model agents. A single static prompt, appended with per-task state at the tail, stays byte-identical across calls and across tasks: the frozen submission prompt served 78% of input tokens from cache (86.6% across its warm tail), against 73% over a three-week development corpus in which prompt edits repeatedly reset the cache. This compounds the few-call design into a small fraction of nominal input compute.

发表机构

  • Proxima Ultra

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑