arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SRHarness:面向智能体符号回归的运行时框架

SRHarness: A Harness for Agentic Symbolic Regression

Zihan Yu, Shixuan Zhou, Hao Huang, Jingtao Ding, Yong Li

arXiv 2609.35501首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SRHarness通过可组合科学动作、持久化状态和轨迹生命周期管理,为智能体符号回归提供结构化运行时支持,在LLM-SRBench上显著提升符号准确率与数值泛化能力。

AI 中文摘要

近期基于智能体的符号回归方法日益依赖大型语言模型来分析数据、选择科学运算,并在长搜索轨迹中细化假设。在这类系统中,性能不仅取决于底层模型和搜索策略,还取决于支持科学搜索的运行时基础设施。我们引入了SRHarness,一个面向智能体符号回归的领域专用运行时框架,其构建围绕三个机制:可组合的科学动作,为原始数据、变换后数据及候选派生量提供统一接口;持久化的科学状态,保留已评估的假设并展示紧凑的面向模型的视图;以及轨迹生命周期管理,协调延续、分支、重启和终止。在LLM-SRBench上,在匹配的LLM主干下,SRHarness持续提升了数值泛化能力和符号恢复性能。使用DeepSeek-v4-flash-0731,SRHarness在LSR-Transform上达到了93.69%的符号准确率,而SR-Scientist为62.16%;在去除科学描述和变量语义的匿名变体上,SRHarness保持了72.97%的准确率,而SR-Scientist为39.64%。在相同的DeepSeek-v4-flash-0731主干下,SRHarness也大幅优于Codex(72.97%对比20.72%),并达到了与使用GPT-5.5的Codex相当的性能,而仅仅为Codex提供相同的科学工具并不能复现这一优势。这些结果表明,有效的智能体符号回归不仅依赖于模型或工具,还依赖于组织科学动作、累积假设和长时程搜索的结构化运行时支持。

英文摘要

Recent agentic symbolic regression approaches increasingly rely on large language models to analyze data, select scientific operations, and refine hypotheses over long search trajectories. In such systems, performance depends not only on the underlying model and search strategy, but also on the runtime infrastructure that supports scientific search. We introduce SRHarness, a domain-specific harness for agentic symbolic regression built around three mechanisms: composable scientific actions that provide a common interface over raw, transformed, and candidate-derived quantities; persistent scientific state that retains evaluated hypotheses and exposes compact model-facing views; and trajectory lifecycle management that coordinates continuation, branching, restart, and termination. On LLM-SRBench, SRHarness consistently improves both numerical generalization and symbolic recovery under matched LLM backbones. With DeepSeek-v4-flash-0731, it achieves 93.69% symbolic accuracy on LSR-Transform, compared with 62.16% for SR-Scientist, and retains 72.97% accuracy on an anonymized variant that removes scientific descriptions and variable semantics, versus 39.64% for SR-Scientist. Under the same DeepSeek-v4-flash-0731 backbone, SRHarness also substantially outperforms Codex (72.97% vs. 20.72%) and reaches performance comparable to Codex with GPT-5.5, while simply providing Codex with the same scientific tools does not reproduce this advantage. These results show that effective agentic symbolic regression depends not only on models or tools, but also on structured runtime support for organizing scientific actions, accumulated hypotheses, and long-horizon search.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑