RankEvolve:一种用于演化排序模型的可靠多智能体自动研究框架
RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models
- Meta
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
RankEvolve提出一种多智能体自动研究框架,通过可执行操作协议和异构编码智能体组合,将执行准确性从45.8%提升至62.5%,并在排序模型演化中实现性能增益。
AI中文摘要:
自动研究智能体,即能够跨迭代提出、实现、训练和评估模型变更的LLM系统,有望自动化应用机器学习的实验循环。在长时间跨度下,执行准确性是一个约束性因素:一次变更可能静默泄漏保留数据、省略归一化、断开梯度或留下未接线的训练/评估标志,从而使昂贵的运行失效并在迭代中累积错误。我们提出了RankEvolve,一个用于演化生成式排序模型的自动研究框架。可执行操作协议(EOP)声明阶段、门控、分支和循环,运行时强制执行编译后的状态机。一个元-元框架将完整的黑盒编码智能体产品(包括Claude Code和Codex)组合为执行图节点,这些节点相互审查和修复彼此的工作。在预算匹配的评估中,异构组合将全知执行准确性从最佳单一产品基线的45.8%提升至62.5%(配对+16.7个百分点,95%置信区间[6.6, 26.7]),同时实现10.4%的静默关键缺陷率。实现的知识层跨迭代携带发现,包括负面结果。在开源HSTU推荐器上的十二次迭代部署中,RankEvolve在MovieLens-20M LARGE上报告了NDCG@10为0.2192(比已发布锚点高+4.48%),在BASE上为0.1948(+2.80%)。由该部署中的事件种子化的ExecML-HSTU为执行准确性评估提供了全知基准。预指定的LitGPT迁移分割复制了推荐之外的异构组合效应(+12.5个百分点,95%置信区间[3.0, 22.0]),配对消融隔离了逐步与全协议指令注入。这些结果刻画了运行时控制的编码智能体产品组合何时能提高执行准确性。
英文摘要:
Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leave a train/eval flag unwired, invalidating expensive runs and compounding error across iterations. We present RankEvolve, an auto-research framework for evolving generative ranking models. An Executable Operating Protocol (EOP) declares phases, gates, branches, and loops, and the runtime enforces the compiled state machine. A meta-meta-harness composes complete black-box coding-agent products, including Claude Code and Codex, as execution-graph nodes that review and repair one another's work. In a budget-matched evaluation, heterogeneous composition raises all-oracle execution accuracy from the best single-product baseline of 45.8 percent to 62.5 percent (paired +16.7 points, 95 percent CI [6.6, 26.7]) while achieving a 10.4 percent silent critical-defect rate. An implemented knowledge layer carries findings, including negative results, across iterations. In a twelve-iteration deployment on the open-source HSTU recommender, RankEvolve reported NDCG@10 of 0.2192 on MovieLens-20M LARGE (+4.48 percent over the published anchor) and 0.1948 on BASE (+2.80 percent). ExecML-HSTU, seeded by incidents from that deployment, provides the oracle benchmark for the execution-accuracy evaluation. A pre-specified LitGPT transfer split replicates the heterogeneous-composition effect beyond recommendation (+12.5 points, 95 percent CI [3.0, 22.0]), and a paired ablation isolates per-step from full-protocol instruction injection. These results characterize when runtime-controlled composition of coding-agent products improves execution accuracy.