ST-Bench:面向科学研究任务的多智能体系统生成的空间-时间基准
ST-Bench: A Spatial-Temporal Benchmark for Multi-Agent System Generation on Scientific Research Tasks
浏览论文内容
中文总结 AI 辅助
ST-Bench基准评估了多智能体系统在科学数据分析任务上的表现,发现多数配置优于单智能体,但优势有条件和成本代价。
中文摘要 AI 辅助
基于LLM的多智能体系统(MAS)的快速进展表明,在编码、数学和问答任务上,它们大幅优于单智能体,这些任务中可执行的测试提供了二元的成功信号。这种优势是否能够迁移到真实的科学数据分析中,仍未得到验证。我们引入了ST-Bench,一个旨在回答两个问题的基准:MAS在复杂的科学数据分析任务上是否优于单智能体,如果是,优势有多大以及额外成本是多少。ST-Bench包含100个数据科学任务,这些任务改编自已发表的涉及水文学、农业和湿地甲烷研究的地球科学研究,扩展为基于额外已发表研究的2067个查询,并由领域专家验证。使用ST-Bench,我们在两种训练协议下评估了五种最新的MAS生成方法,并与基于相同GPT-5骨干的单智能体基线进行比较。十种MAS配置中有九种超过了最便宜的单智能体基线,最强的配置综合得分接近其三倍。这一增益主要归因于覆盖率:训练的工作流在更大比例的查询上产生了现实的数值指标,而在产生现实输出的条件下,这些指标的质量与单智能体基线相当。最强的配置所需的推理时间约为单智能体的四倍,而一个更经济的工作流以不到两倍的成本捕获了大部分收益。MAS的专门化在科学数据分析上带来了可衡量的收益,但这种收益是有条件的,而非普遍的。
英文摘要
The rapid progress of LLM-based multi-agent systems (MAS) has shown that they largely outperform single agents on coding, math, and QA tasks, where executable tests provide a binary success signal. Whether this advantage transfers to real scientific data analysis remains untested. We introduce ST-Bench, a benchmark designed to answer two questions: whether MAS outperform single agents on complex scientific data analysis tasks, and if so, by how much and at what additional cost. ST-Bench contains 100 data science tasks adapted from published Earth science studies across hydrology, agriculture, and wetland methane research, expanded into 2,067 queries grounded in additional published studies and validated by domain experts. Using ST-Bench, we evaluate five recent MAS generation methods under two training protocols, against single-agent baselines on the same GPT-5 backbone. Nine of the ten MAS configurations exceed the cheapest single-agent baseline, with the strongest reaching nearly three times its composite score. This gain is primarily attributable to coverage: trained workflows produce realistic numerical metrics on a larger fraction of queries, while the quality of those metrics, conditional on producing realistic output, is comparable to that of the single-agent baseline. The strongest configuration requires approximately four times the single-agent inference time, whereas a more economical workflow captures the majority of the benefit at less than twice the cost. MAS specialization confers measurable benefit on scientific data analysis, but the benefit is conditional rather than universal.
发表机构
- Rutgers University(罗格斯大学)
- University of Pittsburgh(匹兹堡大学)
- University of Minnesota(明尼苏达大学)
- Oak Ridge National Laboratory(橡树岭国家实验室)
- NEC Labs America(NEC美国实验室)
- University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。