arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25068cs.AI

推荐系统应多久调用一次大语言模型?价值加权路由、监控与季节性稳健性

How Often Should a Recommender Call an LLM? Value-Weighted Routing, Monitoring, and Seasonal Robustness

Bhavtosh Rath

首次发表
浏览论文内容

中文总结 AI 辅助

研究探讨推荐系统调用LLM的频率,提出价值路由器,通过合成模拟分三阶段研究,比较不同路由方式,揭示失败模式,模拟需求激增,得出成本感知路由系统设计原则。

中文摘要 AI 辅助

在廉价启发式方法和昂贵大语言模型(LLM)之间的路由决策通常被视为难题:将难题发送到昂贵路径。但我们认为这种框架不完整,因为难度和商业价值是不同维度,一个难的廉价商品和一个难的昂贵商品错误成本不同。我们提出价值路由器,它是零售商品管道的完全合成模拟,仅使用估计难度和估计价值进行路由,而非真实值。研究分三个阶段。首先,在类别数量与价值呈负相关的合成目录上,将价值加权阈值路由器与仅基于难度和随机基线进行比较。价值加权在召回率与仅基于难度的基线匹配(60%)的同时,精度大幅提高(98.3%对94.3%)。其次,决策记录器和监控器揭示了聚合指标隐藏的失败模式,表明聚合结果几乎完全由类别间差异而非单个项目的区分驱动。第三,模拟黑色星期五需求激增(数量增加2.5倍且向高价值类别转移),比较静态路由器、季节性调整路由器和两种慢路径预算策略。所有结果均来自有实验者定义真实值的受控合成模拟,说明了成本感知路由系统的设计原则,而非经过验证的现实世界声明。

英文摘要

Routing decisions between a cheap heuristic and an expensive large language model (LLM) are typically framed as a difficulty problem: send the hard cases to the expensive path. We argue this framing is incomplete because difficulty and business value are distinct axes - a difficult cheap item and a difficult costly item do not have the same cost of error. We present Value Router, a fully synthetic simulation of a retail merchandising pipeline that routes items using only estimated difficulty and estimated value, never ground truth. The study has three stages. First, a value-weighted threshold router is compared with a difficulty-only and a random baseline on a synthetic catalog with an inverse correlation between category volume and value. Value-weighting matches the difficulty-only baseline's recall of true high-value items (60%) while achieving substantially higher precision (98.3% vs. 94.3%). Second, a decision logger and monitor expose a failure mode hidden by aggregate metrics showing that the aggregate result is driven almost entirely by between-category differences rather than per-item discrimination. Third, a simulated Black Friday demand surge (2.5 volume with a shift toward higher-value categories) compares a static router, a seasonally tuned router, and two slow-path budget policies. All results are from a controlled synthetic simulation with experimenter-defined ground truth and illustrate design principles for cost-aware routing systems rather than validated real-world claims.

↑