发表机构
Meta(Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出人群进化(PE)框架,通过共享实验证据连接局部搜索,实现跨任务迁移和计算密集型场景下的模型发现,在多个任务上显著提升性能。
AI 中文摘要
LLM驱动的进化使得迭代式模型开发成为可能,但两个实际目标仍未得到充分探索:找到能够跨相关任务迁移的模型设计,以及在训练成本高昂时维持改进。我们提出了人群进化(PE),一个协作式、分层的框架,通过共享的实验证据连接正在进行的局部搜索。PE在相关的训练实例上评估代码变更,并共享结果以指导后续提议和向更大训练规模的晋升。对于昂贵的训练目标,PE在小的训练子集上进行搜索,并在全目标训练之前通过同伴和中间评估筛选候选。我们引入了RMD-Bench来评估这两种设置,涵盖排序、观看时间预测、强化学习算法发现以及LLM/VLM预训练。与匹配源迭代下的独立进化相比,PE在排序中将平均最佳局部增益从7.01%提升至8.97%,在观看时间预测中从2.84%提升至3.85%,同时在所有三个联合发现家族中改善了最佳更大规模结果。在观看时间发现中,PE在五个工具中的四个和所有四个提议者上改善了最佳更大规模增益。在共享目标侧校准的新推荐数据集上,每个评估的PE设计在平均性能上均优于参考。在匹配的总GPU计算量下,完成的LLM发现运行产生了2.48%的最佳相对准确率增益和13个成功候选,而直接进化仅为0.92%且无成功候选。在匹配的总GPU计算量下,VLM损失减少达到8.78%,而直接进化为5.05%。
英文摘要
LLM-driven evolution enables iterative model development, but two practical goals remain underexplored: finding model designs that transfer across related tasks and sustaining improvement when training is expensive. We introduce Population Evolution (PE), a collaborative, hierarchical framework that connects ongoing local searches through shared experimental evidence. PE evaluates code changes across related training instances and shares the results to guide subsequent proposals and promotion to larger training scales. For expensive targets, PE searches small training subsets and screens candidates through peer and intermediate evaluations before full-target training. We introduce RMD-Bench to evaluate both settings across ranking, watch-time prediction, RL algorithm discovery, and LLM/VLM pretraining. Compared with standalone evolution at matched source iterations, PE raises mean best local gains from 7.01% to 8.97% in ranking and from 2.84% to 3.85% in watch-time, while improving the best larger-scale outcome in all three joint-discovery families. In watch-time discovery, PE improves best larger-scale gains with four of five harnesses and all four proposers. On new recommendation datasets under shared target-side calibration, every evaluated PE design improves over the reference in mean performance. Under matched total GPU compute, completed LLM discovery runs yield a best relative accuracy gain of 2.48% and 13 successful candidates for PE, versus 0.92% and none for direct evolution. VLM loss reduction reaches 8.78% versus 5.05% under matched total GPU compute.
Comments38 pages, 6 figures, 34 tables