发表机构
University of Birmingham(伯明翰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在推荐系统异构代理选择中,RouteRec框架在成本约束下比较请求级硬选与项目级学习聚合,发现在MovieLens-1M数据集上,请求级选择粗糙,项目级聚合更有前景,不同聚合方式有不同效果。
AI 中文摘要
推荐系统越来越多地面临在异构代理(协同过滤器、序列模型、基于内容的检索器和基于大语言模型的重排器)之间进行选择的问题,然而没有一个代理是始终最佳的。我们使用RouteRec框架,在成本约束下将此选择研究为任务感知代理排序,该框架在四个传统推荐代理和一个大语言模型重排代理上比较请求级硬选择与项目级学习聚合。在MovieLens-1M数据集上,全质量预言机有很大提升空间(HR@10 = 0.584)。在无泄漏的5折交叉验证协议下,硬选择仍低于BM25(0.223对0.254),选择性大语言模型升级也未改善。对于学习聚合,相同协议产生不同结果:仅低成本变体在HR上与BM25匹配且NDCG点估计更高(0.123对0.114),门控全代理聚合在70.2%大语言模型调用下达到HR@10 = 0.295。结论是请求级选择一个完整代理列表在这种稀疏固定候选设置中过于粗糙,项目级聚合是更有前景的行动空间。
英文摘要
Recommender systems increasingly face a choice among heterogeneous agents -- collaborative filters, sequential models, content-based retrievers, and LLM-based rerankers -- yet no single agent is uniformly best. We study this choice as task-aware agent ranking under cost constraints using RouteRec, a framework that compares request-level hard selection with item-level learned aggregation over four traditional recommender agents and one LLM reranker agent. On MovieLens-1M, the full quality oracle has substantial headroom (HR@10 = 0.584), confirming that useful cross-agent signal exists. Under a leakage-free 5-fold out-of-fold protocol, however, hard selection remains below BM25 (0.223 vs. 0.254), and selective LLM escalation does not improve it. The same protocol yields a different outcome for learned aggregation: its cheap-only variant matches BM25 in HR and has a higher NDCG point estimate (0.123 vs. 0.114), while gated all-agent aggregation reaches HR@10 = 0.295 with 70.2\% LLM calls. The resulting lesson is not that routing is solved, but that request-level selection of one complete agent list is too coarse for this sparse fixed-candidate setting; item-level aggregation is the more promising action space.
Comments8 pages, 7 figures. Accepted at AgentSearch 2026 (The First Workshop on Indexing, Retrieval, and Ranking of AI Agents), co-located with SIGIR 2026