arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03177cs.LG

前沿大语言模型是高效的批量优化器:评估连续与离散场景下的推理模型

Frontier LLMs are effective batch optimizers: Assessing reasoning models in continuous and discrete settings

  • Prescient Design, Genentech(基因泰克公司的Prescient Design部门)

机构由 AI 辅助整理,请以论文原文为准。

Frank Hu, Shriram Chennakesavalu, David Graff

AI总结:

该研究评估前沿LLMs在连续与离散场景下作为批量优化器的性能,发现其在数值测试函数上具竞争力但性能较脆弱,在语义丰富场景中表现显著更优,凸显其在特定离散空间优化的有效性。

AI中文摘要:

前沿大语言模型(LLMs)因大规模预训练使其能应对各类优化场景,已成为优化领域颇具吸引力的先验模型。然而,现代推理型LLMs在批量优化场景中的有效性仍未得到充分探索。本文研究当前一代前沿LLMs作为批量优化器在连续与离散场景下的性能。我们发现,LLMs作为数值测试函数的零样本批量优化器具备竞争力,但与经典非LLM优化方法相比,其性能较脆弱;不过,在语义丰富场景中,LLM先验模型的表现显著更优,这表明当处理与预训练数据结构最相似的离散空间时,它们的批量优化行为极为有效。

英文摘要:

Frontier large language models (LLMs) have become attractive priors for optimization due to their large-scale pretraining that enables them to navigate a variety of optimization settings. However, the effectiveness of modern reasoning LLMs in batch optimization settings remains underexplored. Here we investigate the performance of the current generation of frontier LLMs as batch optimizers in both continuous and discrete settings. We find that while LLMs are competitive zero-shot batch optimizers for numerical test functions, their performance is brittle compared to classical non-LLM optimization approaches. However, LLM priors are significantly better in semantically rich settings, indicating that their batch optimization behavior is highly effective when navigating and reasoning over the discrete spaces most similar in structure to their pretraining data.

↑