arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05363cs.LG

全局蒸馏,局部适配:用于可升级推荐的推理蒸馏与乘积类型测试时训练

Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation

Siliang Liu, Mohammad Ghasemi, Sapan Patel, Amin Banitalebi-Dehkordi

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出两级框架,将LLM推理蒸馏为高效学生模型,结合乘积类型测试时训练优化,在可升级推荐任务上实现精度提升,且推理速度远快于直接LLM推理、成本大幅降低。

中文摘要 AI 辅助

可升级推荐旨在识别既能保留顾客购买意图、又能提供升级福利的更高质量替代产品。大语言模型(LLM)可对这类差异进行推理,但直接将其应用于数亿产品对时,在操作层面并不现实。本文提出一种两级框架,将LLM的推理能力蒸馏为高效的非生成式学生模型,并使其决策边界适配特定产品类型的可升级标准。在第一级,检索增强的少样本LLM教师生成结构化关系标签与自然语言原理,这些原理通过对齐和对比目标监督紧凑的嵌入对分类器;推理阶段,学生模型仅使用两个预计算的768维产品嵌入,无需调用LLM或生成文本。在包含8352对的固定人工标注基准上,参数规模为1550万的四类推理蒸馏学生模型的AUC达到0.924(95%置信区间[0.918, 0.929]),而仅用标签训练的四类学生模型AUC为0.912。在第二级,乘积类型测试时训练(PT-TTT)利用少样本演示,在冻结的学生模型上优化轻量级特定类别适配器;PT-TTT使AUC从0.924提升至0.941,平均精度从0.920提升至0.940。在包含10万对的代理目录上,该蒸馏学生模型在单台8-GPU机器上的推理速度约为直接LLM推理的5000倍,估计成本则低约10000倍。

英文摘要

Trade-up recommendation identifies higher-quality alternatives that preserve a customer's purchase intent while offering upgraded benefits. Large language models (LLMs) can reason about such distinctions, but applying them directly to hundreds of millions of product pairs is operationally impractical. We introduce a two-level framework that distills LLM reasoning into an efficient non-generative student and adapts its decision boundary to product-type-specific trade-up criteria. At Level 1, a retrieval-augmented few-shot LLM teacher generates structured relation labels and natural-language rationales. These rationales supervise a compact embedding-pair classifier through alignment and contrastive objectives; at inference, the student uses only two precomputed 768-dimensional product embeddings, with no LLM calls or text generation. On a fixed human-annotated benchmark of 8,352 pairs, a 15.5M-parameter four-class reasoning-distilled student achieves AUC 0.924 (95% CI [0.918, 0.929]), compared with 0.912 for the four-class label-only student. At Level 2, product-type test-time training (PT-TTT) uses few-shot demonstrations to optimize lightweight category-specific adapters over the frozen student. PT-TTT improves AUC from 0.924 to 0.941 and average precision from 0.920 to 0.940. On a 100K-pair proxy catalog, the distilled student on a single eight-GPU machine is approximately 5,000x faster and 10,000x lower in estimated cost than direct LLM inference.

发表机构

  • Amazon Everyday Essentials Technologies(亚马逊日常必需品技术公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑