arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgentHPOBench:用于评估LLM智能体作为序列超参数优化器的基准

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

Tianyu Huai, Tingshuo Fan, Xinchi Chen, Yining Zheng, Yuxin Wang, Shuang Chen, Jie Zhou, Xuanjing Huang

arXiv 2607.29626首次发表:更新:

AI 中文总结

AgentHPOBench是含30个任务的序列基准,用于评估LLM智能体的超参数优化能力,实验显示其在多领域有优化能力,但在迭代优化等方面存在局限。

AI 中文摘要

随着大型语言模型(LLMs)从代码补全系统演变为自主科学智能体,评估其开展实验的能力愈发重要。现有基准通常聚焦静态代码生成、论文复现或最终答案正确性,却未直接评估智能体能否解读实验证据并以此指导后续超参数决策。为填补这一空白,我们推出AgentHPOBench,这是一个包含7个研究类别下30个可执行机器学习任务的序列基准。每个任务以经验证的基线运行开始,之后智能体执行若干序列干预。每一步中,智能体观察累积的配置、指标和日志,再提出下一个有效配置。我们在统一协议下评估12种广泛使用的智能体及传统超参数优化(HPO)基线。结果表明,当前智能体在各领域展现出可测量的实验优化能力,但在持续迭代优化、复杂日志诊断以及向报告参考性能稳步推进方面仍存在明显局限。

英文摘要

As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmarks typically focus on static code generation, paper replication, or final answer correctness, but do not directly assess whether agents can interpret experimental evidence and use it to guide subsequent hyperparameter decisions. To address this gap, we introduce AgentHPOBench, a sequential benchmark comprising 30 executable machine learning tasks across seven research categories. Each task begins with a validated baseline run, after which an agent performs several sequential interventions. At each step, the agent observes the accumulated configurations, metrics, and logs before proposing the next valid configuration. We evaluate 12 widely used agents and conventional HPO baselines under a unified protocol. The results show that current agents exhibit measurable experimental optimization ability across domains, but still face clear limitations in sustained iterative refinement, complex log diagnosis, and consistent progress toward reported reference performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑