发表机构
ByteDance; School of Computing and Data Science, The University of Hong Kong; Huazhong University of Science and Technology; Shenzhen Loop Area Institute; Hong Kong University of Science and Technology (Guangzhou); School of Information Technology, Carleton University(字节跳动; 香港大学计算与数据科学学院; 华中科技大学; 深圳河套学院; 香港科技大学(广州); 卡尔顿大学信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出OptiCom统一框架,通过状态条件组合机制,在共享配置空间中动态协调LLM驱动的优化过程,显著提升优化性能,在32个基准组中平均排名1.72。
AI 中文摘要
大型语言模型(LLMs)越来越多地被部署用于通过迭代优化来解决复杂的科学和实际问题。然而,随着候选质量、失败模式和资源预算的演变,动态协调多样化的搜索机制仍然是一个关键的开挑战。针对性的经验诊断表明,机制的有效性高度依赖于状态。受此启发,我们分析了单个决策如何驱动最终结果,将共享预算下的预期最终改进分解为累积决策机会减去累积选择损失。在此机会-损失理论基础指导下,我们提出了OptiCom,一个统一的框架,在共享配置空间C=(A,Q,O,E,M,S)中表示LLM驱动的优化器,分别对应工件、查询、操作符、评估、记忆和策略。在此空间内运行,一个快速的基于LLM的优化控制器通过结构化的动作包动态组合即时机制,而一个较慢的策略适配器则根据累积的轨迹反馈优化长期选择偏好、操作符权重和模板。在32个基准组上的全面评估证明了该框架的优越性:OptiCom在14个评估配置中实现了平均最大分数排名1.72,并在23个组中获得了最高分。最终,这些结果凸显了OptiCom作为稳健的LLM测试时扩展的通用范式的广泛适用性和高可扩展性。
英文摘要
Large language models (LLMs) are increasingly deployed to solve complex scientific and practical problems via iterative optimization. However, dynamically coordinating diverse search mechanisms as candidate quality, failure modes, and resource budgets evolve remains a critical open challenge. Targeted empirical diagnostics reveal that mechanism effectiveness is highly state-dependent. Motivated by this, we analyze how individual decisions drive final outcomes, decomposing the expected terminal improvement under a shared budget into cumulative decision opportunities minus cumulative selection losses. Guided by this opportunity-loss theoretical foundation, we propose OptiCom, a unified framework that represents LLM-driven optimizers within a shared configuration space: C=(A,Q,O,E,M,S), corresponding to artifact, query, operator, evaluation, memory, and strategy. Operating within this space, a fast LLM-based Optimization Controller dynamically composes immediate mechanisms through structured Action Packages, while a slower Strategy Adapter refines long-term selection preferences, operator weights, and templates based on accumulated trajectory feedback. Comprehensive evaluations across 32 benchmark groups demonstrate the superiority of framework: OptiCom achieves an average Max-score rank of 1.72 among 14 evaluated configurations, securing the top score in 23 groups. Ultimately, these results highlight the broad applicability and high extensibility of OptiCom as a general-purpose paradigm for robust LLM test-time scaling.
Comments43 pages, 11 figures