arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

优化低成本部署强模型:面向进化优化的成本感知跨层迁移

Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization

Tal Oved, Roi Pony, Oshri Naparstek, Udi barzelay

arXiv 2608.10694首次发表:更新:

发表机构

IBM Research(IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出成本感知跨层迁移方法,解耦LLM角色,在低成本层级完成大部分搜索,跨层部署提示词,在多任务多模型上实现成本大幅降低且性能不劣于同层级优化。

AI 中文摘要

针对大语言模型(LLM)提示词与智能体程序(如GEPA)的进化优化,核心成本来自适应度评估:对每个候选方案打分需要让回答LLM在验证集上运行,因此评估器的价格层级决定了总搜索成本。我们通过解耦LLM的三类角色来重构搜索流程:将高吞吐量的回答角色放在最便宜的层级,预留强模型用于罕见的反思/变异算子,再利用向上跨层迁移将低成本进化得到的提示词部署到更强的目标模型上。我们提出了成本可控的刻画方法,明确低成本层级搜索何时可替代目标层级搜索、何时会失效。在四个任务(HotpotQA、IFBench、LiveBench-Math、HoVer)和四个模型家族的十一个模型上,该方法得到的提示词与同层级优化的表现相当或更优,同时将超过96%的搜索令牌放在最便宜层级,搜索成本降低5.6至14倍;若推理层级在每次适应度调用时生成长思维链,成本可降低25至54倍。

英文摘要

Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs an answering LLM over a validation set, so the evaluator's price tier dictates total search cost. We restructure that search by decoupling the three roles an LLM plays, running the high-volume answering role on the cheapest tier, reserving a strong model for the rare reflection/variation operator, then exploiting upward cross-tier transfer to deploy the cheaply evolved prompt on a stronger target. We contribute a cost-controlled characterization of when cheap-tier search substitutes for target-tier search, and where it fails. Across four tasks (HotpotQA, IFBench, LiveBench-Math, HoVer) and eleven models in four model families, the resulting prompt matches or exceeds same-tier optimization while placing over 96% of search tokens on the cheapest tier, at 5.6-14x lower search cost, rising to 25-54x where reasoning tiers emit long chains of thought on every fitness call.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑