发表机构
University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EvoRank利用LLM引导的进化循环自动发现多目标学习排序流水线,在Expedia数据集上以低成本超越调优基线,并揭示收益多为适应度噪声,提出可预先测量的余量门来预测进化是否有效。
AI 中文摘要
我们提出了EvoRank,一个开放自主的排序工程师:一个由LLM引导的进化循环,能够为多目标电子商务搜索发现完整的学习排序流水线(特征、模型、损失函数、集成)。在Expedia ICDM 2013数据集上,以相关性、转化率和收入作为竞争目标,三次独立运行均在50次迭代内(约十美元成本)收敛到可解释的流水线,这些流水线在6万个保留查询上击败了Optuna调优的LambdaMART,该优势在全数据规模下依然存在,并跻身原始竞赛前6%。首个仅进化训练目标的活动构建了核心设计规则:它在其选择折(用于挑选优胜者的小数据集)上似乎有效,而转移审计(在保留数据上重新评分优胜者)显示这些收益几乎完全是适应度噪声(其自身评分的随机性),且无论是注入领域知识还是更丰富的诊断反馈,都未改变转移结果。决定性的数量可提前测量:相对于适应度噪声的搜索空间余量。我们将其封装为一个余量门,在任何LLM支出之前预测循环是否会有回报,并发布了系统、审计工具以及带有防护措施的失败模式目录,以便团队将这一流程应用于自己的排序栈。
英文摘要
We present EvoRank, an open autonomous ranking engineer: an LLM-guided evolutionary loop that discovers complete Learning-to-Rank pipelines (features, models, losses, ensembles) for multi-objective e-commerce search. On the Expedia ICDM 2013 dataset, with relevance, conversion, and revenue as competing objectives, three independent runs each converge within 50 iterations (about ten dollars) on interpretable pipelines that beat an Optuna-tuned LambdaMART on 60k held-out queries, an advantage that persists at full data scale and places in the top 6 percent of the original competition. A first campaign, evolving only training objectives, builds the central design rule: it appeared to work on its selection fold (the small dataset it uses to pick winners) while a transfer audit, re-scoring winners on held-out data, showed the gains were almost entirely fitness noise (the randomness of its own scoring), and neither seeded domain knowledge nor richer diagnostic feedback changed what transferred. The deciding quantity is measurable in advance: search-space headroom relative to fitness noise. We package this as a headroom gate that predicts, before any LLM spend, whether the loop will pay off, and we release the system, the auditing tools, and a catalog of failure modes with their guardrails, so teams can apply the procedure to their own ranking stacks.