arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

差异推理路由:在电商中实现成本感知的大语言模型标注

The Differential Reasoning Router: Operationalizing Cost-Aware LLM Annotation in E-commerce

Cheng Lyu, Jingyue Zhang, Vinny DeGenova, Mengwei Li, Yuanli Pei

arXiv 2608.30224首次发表:更新:

发表机构

Wayfair(Wayfair)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对电商LLM标注冷启动问题,提出DRR框架,通过联合优化模型选择与人工升级实现自适应路由,在保持准确率的同时节约60%以上推理令牌成本。

AI 中文摘要

大语言模型(LLM)正越来越多地用于电商中结构化产品数据的标注,但早期部署常面临冷启动问题:仅能获取有限的预发布标签,昂贵推理的价值尚不明确,且需人工审核才能让系统获得大规模信任。这一挑战在基于规则的标注工作流中尤为常见,其中每个条目必须满足多条业务规则,模型错误与模糊的规则边界都会影响最终决策。我们提出差异推理路由(Differential Reasoning Router,DRR),这是一种用于冷启动LLM标注的成本感知框架,可联合优化模型选择与人工升级流程。DRR不将推理模型视为默认后备,而是在样本和业务规则层面分别估计直接模型与推理模型的成功概率,从而实现自适应路由:简单案例由直接模型处理,推理仅用于预期能改进决策的案例,而可能出现双重失败或规则分歧的案例则升级至人工标注者。生成的标签为提示工程、监督微调、校准及规则优化提供了针对性的真值,支持从人工密集型冷启动标注逐步转向高置信度自动路由。在某电商生产工作流中,DRR达到了最强的基于置信度的路由的准确率对等,同时实现了超过60%的推理令牌成本节约。

英文摘要

Large Language Models (LLMs) are increasingly used to annotate structured product data in e-commerce, but early deployment often begins as a cold-start problem: only limited pre-launch labels are available, the value of expensive reasoning is unknown, and human review is needed before the system can be trusted at scale. This challenge is especially common in rule-based annotation workflows, where each item must satisfy multiple business rules and both model errors and ambiguous rule boundaries affect final decisions. We introduce the Differential Reasoning Router (DRR), a cost-aware framework for cold-start LLM annotation that jointly optimizes model selection and human escalation. Rather than treating a reasoning model as a default fallback, DRR estimates separate success probabilities for a direct model and a reasoning model at both the sample and business-rule levels, enabling adaptive routing: easy cases are handled directly, reasoning is reserved for cases where it is expected to improve the decision, and likely double-failure or rule-disagreement cases are escalated to human annotators. The resulting labels provide targeted ground truth for prompt engineering, supervised fine-tuning, calibration, and rule refinement, enabling a gradual shift from human-heavy cold-start annotation toward high-confidence automated routing. In a production e-commerce workflow, DRR reaches accuracy parity with the strongest confidence-based router while achieving more than 60\% reasoning-token cost savings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑