RecGPT-Mobile-V2 技术报告
RecGPT-Mobile-V2 Technical Report
浏览论文内容
中文总结 AI 辅助
本文提出 RecGPT-Mobile-V2 端到端框架,解决个性化查询预测的设备端挑战,经实验验证其在查询质量、失败率等指标上表现更优,推理策略更具针对性。
中文摘要 AI 辅助
个性化查询预测将隐含的行为信号(点击、收藏、购买及购买后探索)映射为明确的检索意图,而设备端部署使该任务极具挑战性:行为轨迹存在噪声且多尺度,单条轨迹可能对应多个有效查询,统一推理策略要么在简单实例上消耗不必要的计算,要么为复杂实例分配的算力不足。本文提出 RecGPT-Mobile-V2,这是一个端到端框架,将意图质量与执行效率作为耦合目标进行分阶段设计。该框架将异构交互转换为保留证据的轨迹,通过领域适应和监督对齐建立推荐原生基础,且仅在分组推演满足基础与效用标准后应用推理成本优化。生成的教师模型被蒸馏为紧凑的学生模型,采用低位执行、结构化压缩及预算感知的设备-云路由部署。在对齐的思维链(CoT) ablation 实验中,聚焦证据的短推理使 ROUGE-L 从 0.228 提升至 0.315,Jaccard 从 0.174 提升至 0.248,同时略优于完整五阶段推理。在受控强化学习(RL)对比中,完整奖励公式将查询质量从仅质量 RL 下的 73.2% 提升至 78.6%,将硬失败率从 3.6% 降至 1.6%,并将中位 CoT 长度从 62 个 token 降至 14 个 token。在线检索分析进一步表明,查询召回通道检索到的商品与现有召回通道呈现的商品互补。总体而言,这些发现支持面向充分性而非统一简短的推理:保留决策相关证据,仅在可能改进预测查询时分配额外计算。
英文摘要
Personalized Query prediction maps implicit behavioral signals---clicks, favorites, purchases, and post-purchase exploration---to explicit retrieval intent. On-device deployment makes this task particularly challenging: behavioral trajectories are noisy and multi-scale, multiple Queries may be valid for a single trajectory, and a uniform reasoning policy either expends unnecessary computation on simple instances or allocates insufficient capacity to complex ones. We introduce RecGPT-Mobile-V2, an end-to-end framework that treats intent quality and execution efficiency as coupled objectives within a staged design. The framework transforms heterogeneous interactions into an evidence-preserving trajectory, establishes a recommendation-native foundation through domain adaptation and supervised alignment, and applies reasoning-cost optimization only after grouped rollouts meet grounding and utility criteria. The resulting teacher is distilled into a compact student deployed with low-bit execution, structured compression, and budget-aware device--cloud routing. In an aligned CoT ablation, an evidence-focused short rationale increases ROUGE-L from 0.228 to 0.315 and Jaccard from 0.174 to 0.248, while slightly outperforming the full five-stage rationale. In the controlled RL comparison, the complete reward formulation improves Query quality from 73.2% under quality-only RL to 78.6%, lowers the hard-failure rate from 3.6% to 1.6%, and reduces the median CoT length from 62 to 14 tokens. Online retrieval analysis further indicates that the Query recall channel retrieves inventory complementary to that surfaced by established recall channels. Collectively, these findings support sufficiency-oriented rather than uniformly short reasoning: retain decision-relevant evidence and allocate additional computation only when it is likely to improve the predicted Query.