arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

规则到工具:面向科学计算中LLM智能体的可执行检查

Rules to Tools: Executable Checks for LLM Agents in Scientific Computing

Jingjie Ning, Guojiang Zhao, Chen Xu, Shanshan Zhong, Xiaochuan Li, Ji Zeng, Guolin Ke

arXiv 2610.00313首次发表:更新:

发表机构

Carnegie Mellon University; DP Technology(卡内基梅隆大学; 深势科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对科学计算中的LLM智能体,提出规则到工具(R2T)方法,将书面要求转化为可执行检查,实验表明在多个SciCode任务队列中提升修复成功率并降低智能体输出成本。

AI 中文摘要

科学编码智能体以书面形式接收方程、边界条件和输出要求,然后必须评估它们所修改的程序。规则到工具(R2T)提供了针对公开科学要求的预制备可执行检查。匹配的SciCode修复组共享书面检查、起始程序、模型和预算;工具组接收可调用的实现。在两个任务ID队列中,使用文本的完整修复为26/30,使用预制备检查的为29/30。三个任务ID偏好工具,一个偏好文本,十一个持平。八ID队列得分为13/16对比15/16,差异的任务簇自助法95%区间为[-12.5, 43.75]个百分点。更大的共享定义SciCode队列每组持平于13/24。五个开发暴露的任务使用替代起始程序得分为3/10对比7/10。工具组偏好任务17、77和11;初始检查标记任务17,并对任务77和11报告无违规。任务37偏好文本且无初始报告违规。一个全新的源到Python分支也达到15/16,与专用命令的汇总匹配。在匹配的PDE比较中,详细文本得分为23/24,检查得分为24/24,检查的报告模型输出低31.2%。智能体侧输出节省因队列而异,而公共CPU使用在两个任务ID队列中均上升。这些结果衡量了任务依赖的修复结果和预制备检查下的智能体侧成本。

英文摘要

Scientific coding agents receive equations, boundary conditions, and output requirements in writing, then must assess the programs they revise. Rules to Tools (R2T) supplies prepared executable checks of public scientific requirements. Matched SciCode repair groups share written checks, starting programs, model, and budgets; the tool group receives a callable implementation. Across two task-ID cohorts, complete repair is 26/30 with text and 29/30 with the prepared checks. Three task IDs favor tools, one favors text, and eleven tie. The eight-ID cohort scores 13/16 versus 15/16, with a task-cluster bootstrap 95% interval of [-12.5, 43.75] percentage points for the difference. The larger shared-definition SciCode cohort ties at 13/24 per group. Five development-exposed tasks with alternate starting programs score 3/10 versus 7/10. The tool group favors tasks 17, 77, and 11; initial checks flag task 17 and report no violation for tasks 77 and 11. Task 37 favors text and has no initial reported violation. A fresh source-through-Python arm also reaches 15/16, matching the dedicated command's aggregate. In a matched PDE comparison, detailed text scores 23/24 and checks score 24/24, with 31.2% lower reported model output for checks. Agent-side output savings vary by cohort, while public CPU use rises in both task-ID cohorts. These results measure task-dependent repair outcomes and agent-side costs with prepared checks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑