发表机构
University of California, Riverside(加利福尼亚大学河滨分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出 ToolGate 流水线,用于构建工具依赖型科学基准,通过三道关卡筛选语言模型生成的任务,在 FEniCSx 上验证后得到 128 个唯一有效任务,减轻了专家的重复工作。
AI 中文摘要
科学基准通常由领域专家构建,他们编写任务并交叉核对彼此的工作,或改编教科书、已发表论文及在线资源中的现有材料。这些途径能产生高质量的评估,但每项任务都需要大量人工。语言模型可通过快速生成候选任务来减少这种重复工作,剩下的问题是如何验收。我们针对的是答案需要使用专业软件计算而非仅靠独立推理的科学问题。若候选任务的脚本运行失败或返回不同答案,则该任务无效;若模型无需使用软件就能回答该任务,则该任务是 trivial 的。我们提出 ToolGate,它将每个生成的任务视为一个提案,仅当通过三道关卡时才保留该任务。第一,可执行解决方案脚本必须在使用科学软件运行时复现所提出的答案;第二,随机无工具筛选会拒绝模型仅从提示就能解决的候选任务;第三,使用工具的智能体必须在固定时间限制内解决每个留存的任务。我们在 FEniCSx 中实例化 ToolGate,进行了 500 次生成尝试。本地验证关卡保留了 478 个候选任务。为最终报告,我们在生成后重新筛选该池:两次随机无工具筛选从报告池中排除了 222 个,使用中等推理(API 默认)的直接 GPT-5.5 API 调用又排除了 121 个。在剩余的 135 个任务中,可访问 FEniCSx 的 GPT-5.5 Codex CLI 智能体解决了 130 个;精确去重后得到 128 个唯一协议留存任务。ToolGate 将重复的答案检查和难度筛选转变为可审计的流程,同时将领域设计和最终审核留给专家。
英文摘要
Scientific benchmarks are commonly built by domain experts who write tasks and cross-check one another's work, or who adapt existing material from textbooks, published papers, and online resources. These routes can produce strong evaluations, but they require substantial per-item labor. Language models can reduce this repeated work by proposing candidates quickly. The remaining problem is acceptance. We target scientific questions whose answers require computations with specialist software rather than unaided reasoning alone. A candidate is invalid if its script fails or returns a different answer, or trivial if a model answers it without the software. We present ToolGate, which treats every generated item as a proposal and keeps it only if three gates pass. First, an executable solution script must reproduce the proposed answer when run with the scientific software. Second, randomized no-tool screening rejects candidates that models can already solve from the prompt alone. Third, a tool-using agent must solve each survivor within a fixed time limit. We instantiate ToolGate in FEniCSx with 500 generation attempts. The local-verification gate retains 478 candidates. For final reporting, we rescreen this pool after generation: two randomized no-tool screens exclude 222 from the reported pool, and direct GPT-5.5 API calls at medium reasoning (the API default) exclude another 121. Of the remaining 135, a GPT-5.5 Codex CLI agent with access to FEniCSx solves 130; exact deduplication leaves 128 unique protocol survivors. ToolGate turns repeated answer checking and difficulty screening into an auditable process while leaving domain design and final review to experts.
Comments7 pages, 2 figures