发表机构
Amazon(亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对工业可解释推荐系统的高成本问题,提出将 LLM 解释生成与选择分离的方案,发现 pairwise 排序方法在离线解释选择中优于单动作强化学习,且构建成本低、延迟小。
AI 中文摘要
基于大语言模型(LLM)构建的工业可解释推荐系统会产生巨大的服务成本:每个请求都会触发一次 LLM 生成,延迟达数百毫秒,成本随流量线性增长。我们将生成与选择分离:解释提前生成为固定候选池(六种提示风格、两款商用 LLM),一个驻留在 CPU 的小型选择器在请求时选择其中一个。该架构无需 GPU,返回时间在 100 毫秒以内。我们的主要基准是包含 2958 对的 XRec Google Local 子集,评估六个离线池选择器(LambdaRank、PPO、GRPO、DPO、师生蒸馏)和三个知识图谱(KG)路径选择器(随机游走、边不相交枚举、MMR 重排序路径)。以包含 300 对的 MovieLens-1M 拆分集(以 Claude-Sonnet-4.5 为参考)作为内部跨数据集验证,因为该设置不存在公开基准。所有变体均采用与 XRec 和 G-Refer 相同的 BERTScore-F1 协议,在五个随机种子上取平均值。LambdaRank 在 Google Local 上达到 F1 = 0.500,超过 G-Refer 和 XRec;在 MovieLens-1M 验证集上达到 F1 = 0.329。由于种子方差低于 0.003 F1,排序可靠: pairwise 学习排序优于单动作强化学习(PPO、GRPO、DPO),后者每次仅使用一个带标签的候选,浪费 K-1 个标签。KG 路径系列目标不同:三个变体在 Google Local 上达到 USR = 1.000,在 MovieLens-1M 上达到 0.997-1.000,因为每个请求的路径 grounding 会为每个查询生成唯一输出,避免影响缓存 LLM 输出的模板崩溃故障。对比 Claude 3 Haiku 和 Claude Haiku 4.5 的生成器池研究显示,F1 偏移很小(0.001-0.006),同时保持选择器排序:选择器和生成器可独立评估,尽管绝对 F1 取决于生成器。在商用硬件上的端到端构建成本接近 15 美元。
英文摘要
Industrial explainable-recommendation systems built on LLMs incur a substantial serving cost: each request triggers an LLM generation, with latency in the hundreds of milliseconds and cost that scales linearly with traffic. We separate generation from selection: explanations are produced ahead of time as a frozen candidate pool (six prompt styles, two commodity LLMs), and a small CPU-resident selector picks one at request time. The stack needs no GPU and returns in under 100 ms. Our primary benchmark is a 2,958-pair XRec Google Local subset, evaluating six offline-pool selectors (LambdaRank, PPO, GRPO, DPO, teacher-student distillation) and three KG-path selectors (random walks, edge-disjoint enumeration, MMR-reranked paths). A 300-pair MovieLens-1M split with Claude-Sonnet-4.5 references serves as an internal cross-dataset check, since no public benchmark exists for this setting. All variants use the same BERTScore-F1 protocol as XRec and G-Refer, averaged across five seeds. LambdaRank reaches F1 = 0.500 on Google Local, exceeding both G-Refer and XRec, and F1 = 0.329 on the MovieLens-1M check. With seed variance below 0.003 F1, the ordering is reliable: pairwise learning-to-rank outperforms single-action RL (PPO, GRPO, DPO), which use only one labelled candidate per rollout, leaving K-1 labels unused. The KG-path family targets a different objective: all three variants reach USR = 1.000 on Google Local and 0.997-1.000 on MovieLens-1M, since per-request path grounding yields a unique output per query, avoiding template-collapse failures affecting cached-LLM outputs. A generator-pool study comparing Claude 3 Haiku and Claude Haiku 4.5 shows small F1 shifts (0.001-0.006) while preserving selector ranking: selector and generator can be evaluated independently, though absolute F1 depends on the generator. End-to-end build cost is near $15 on commodity hardware.
CommentsThis is an extended version of a 3-page paper accepted to the RecSys 2026 Research and Practice Notes track