WMLLM:基于“预测后行动”世界建模的自进化优化智能体
WMLLM: Self-Evolving Optimization Agents via Predict-Then-Act World Modeling
浏览论文内容
中文总结 AI 辅助
WMLLM是基于“预测后行动”世界建模的自进化优化智能体框架,结合多轮优化、种群搜索与强化学习,在有限预算下的多目标分子优化等黑盒任务中实现了最优样本效率与性能。
中文摘要 AI 辅助
黑盒优化问题因搜索空间庞大、结构薄弱且维度高而极具挑战性。现有方法常因依赖直接候选生成或试错式优化而样本效率低下。提升搜索效率的自然途径是采用世界建模,其可在高成本评估前识别有前景的优化方向。大型语言模型凭借其隐含知识,能以可观的精度预测这些候选方案的结果。受此观察启发,我们提出WMLLM,这是一种基于“预测后行动”世界建模的自进化优化智能体框架。该智能体先预测有前景的方向,再行动生成候选方案。结合智能体多轮优化、基于种群的搜索及强化学习,WMLLM在搜索过程中同时优化其隐含世界模型和优化策略。在黑盒优化任务(尤其是多目标分子优化)上的实验表明,WMLLM提升了样本效率和最终优化性能。在多目标分子优化基准测试中,WMLLM在有限评估预算下取得了当前最优结果。
英文摘要
Black-box optimization problems remain challenging because of large, weakly structured, and high-dimensional search spaces. Existing methods often suffer from poor sample efficiency because they rely on direct candidate generation or trial-and-error refinement. A natural way to improve search efficiency is to use world modeling, which can help identify promising optimization directions before costly evaluation. Large language models can predict the outcomes of these candidates with nontrivial accuracy because of their implicit knowledge. Motivated by this observation, we propose WMLLM, a self-evolving optimization-agent framework based on predict-then-act world modeling. The agent first predicts promising directions and then acts to generate candidates. Combined with agentic multi-turn refinement, population-based search, and reinforcement learning, WMLLM refines both its implicit world model and its optimization strategy during search. Experiments on black-box optimization tasks, especially multi-objective molecular optimization, show that WMLLM improves sample efficiency and final optimization performance. On the multi-objective molecular optimization benchmark, WMLLM achieves state-of-the-art results under a limited evaluation budget.
发表机构
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
- Zhongguancun Academy(中关村学院)
机构由 AI 辅助整理,请以论文原文为准。