发表机构
Kuaishou Technology(快手科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RecHarness是一种多臂老虎机路由智能体框架,用于自动化优化推荐系统模型,通过分离方向选择与假设生成、引入跳盆机制提升稳定性,在多任务测试及在线A/B测试中均实现性能提升。
AI 中文摘要
优化现代推荐系统模型仍高度依赖工程师手动迭代架构、目标函数和训练策略的变更。虽然基于大语言模型(LLM)的智能体可自动化该试错过程,但让LLM同时选择修改方向并生成具体假设往往会在有限实验预算下导致搜索不稳定。受上述挑战启发,我们提出RecHarness,一种用于自动化推荐系统模型优化的多臂老虎机路由智能体框架。RecHarness将优化过程分为两步:多臂老虎机路由器根据历史验证反馈选择下一个修改方向,LLM在选定方向内生成具体优化假设和可执行代码编辑。为维持长周期探索,当局部编辑陷入停滞时,RecHarness使用跳盆机制激活结构跳臂。在多个推荐任务、数据集和模型主干上,RecHarness比LLM推理搜索实现了更稳定的性能提升,且更高效地利用有限试验预算。在某大型短视频广告平台为期7天的在线A/B测试中,选定候选方案使ADVV提升2.084%、营收提升0.534%、曝光量提升0.559%。代码可在该URL获取。
英文摘要
Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-strategy changes. While LLM-based agents can automate this trial-and-error process, allowing the LLM to both select modification directions and generate concrete hypotheses often leads to unstable search under limited experiment budgets. Inspired by the above challenge, we propose RecHarness, a Bandit-Routed Agentic Harness for automated recommender model optimization. RecHarness separates the optimization process into two steps: a bandit router selects the next modification direction according to historical validation feedback, while the LLM generates a concrete optimization hypothesis and executable code edit within the selected direction. To sustain long-horizon exploration, RecHarness uses a jump-basin mechanism to activate a structural-jump arm when local edits stagnate. Across multiple recommendation tasks, datasets, and model backbones, RecHarness achieves more stable performance improvements and uses limited trial budgets more effectively than LLM-reasoning search. During a 7-day online A/B test on a large-scale short-video advertising platform, the selected candidate improves ADVV by 2.084%, Revenue by 0.534%, and Exposure by 0.559%. Code is available at https://github.com/6lyc/RecHarness.
Comments9 pages, 2 figures