FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics
FML-bench:从搜索动力学视角对AI研究代理策略的受控研究
机构 * National University of Singapore(国立新加坡大学) ; Tsinghua University(清华大学) ; University of Minnesota(明尼苏达大学) ; Weco ; Meta
专题命中 工具调用 :agent(title,abstract);分类 cs.AI、cs.LG
AI总结 本文提出FML-Bench基准,通过分离策略与基础设施并定义过程级指标,评估六种代理策略,发现贪婪爬山法接近最优树搜索,且自适应策略基于搜索密度切换可超越其他代理。
Comments Our benchmark is available at: https://github.com/qrzou/FML-bench