arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01095cs.SEhep-ex

基于图基础软件知识的高能物理实验中可靠的大语言模型生成程序

Reliable LLM-Generated Programs for High-Energy Physics Experiments through Graph-Grounded Software Knowledge

Yue Sun, Tong Liu, Yipu Liao, Jingde Chen, Ke Li

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLM生成高能物理实验程序不可靠问题,提出结合软件知识图检索等的基础系统,在ROOT基准测试中显著提升程序执行成功率,成本增幅小且可迁移至其他框架。

中文摘要 AI 辅助

从现代粒子物理实验中提取物理信息需要在大型且高度互联的软件生态系统之上实现多阶段分析。通用大语言模型(LLM)针对此类任务常生成不可靠的程序,因为仅用户请求很少能明确所需的API、依赖项和使用规范。我们在生成前梳理这些软件关系,并在推理时检索与任务相关的知识。以开源ROOT框架作为代表性可复现测试平台,我们评估了一套完整的基础系统,该系统结合了异构软件知识图的混合检索、技能筛选的工作流示例以及执行引导修复。在275个ROOT任务的基准测试中,在Claude Code编排下,基础系统使首次尝试执行成功率从58.5%提升至76.0%;在独立编排下,该成功率从51.3%提升至64.0%。最终成功完成率分别从90.5%提升至96.0%和从78.9%提升至90.9%,而每个成功任务的平均生成成本仅分别增加1.3%和3.2%。该增益在强编码智能体下依然存在,表明即使已存在智能体脚手架,显式软件知识仍具价值。由于该方法捕获的是大型代码库共有的软件关系,而非ROOT或特定模型特有的事实,因此它应能迁移至其他实验框架和专有软件,尤其适用于文档稀疏或内部依赖复杂的场景。

英文摘要

Extracting physics information from modern particle-physics experiments requires multistage analyses implemented on top of large and highly interconnected software ecosystems. General-purpose large language models (LLMs) often produce unreliable programs for such tasks because a user request alone rarely specifies the required APIs, dependencies, and usage conventions. We organize these software relations before generation and retrieve task-relevant knowledge at inference time. Using the open-source ROOT framework as a representative and reproducible testbed, we evaluate a complete grounding system that combines hybrid retrieval over a heterogeneous software knowledge graph, skill-selected workflow examples, and execution-guided repair. On a benchmark of 275 ROOT tasks, grounding improves first-attempt execution from 58.5% to 76.0% under Claude Code orchestration and from 51.3% to 64.0% under standalone orchestration. Final success increases from 90.5% to 96.0% and from 78.9% to 90.9%, respectively, while the average generation cost per successful task increases by only 1.3% and 3.2%. The gains persist under a strong coding agent, indicating that explicit software knowledge remains valuable even when agentic scaffolding is already in place. Because the method captures software relations common to large codebases rather than facts specific to ROOT or a particular model, it should transfer to other experiment frameworks and proprietary software, especially where documentation is sparse or internal dependencies are complex.

发表机构

  • Institute of High Energy Physics, Chinese Academy of Sciences(中国科学院高能物理研究所)
  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑