arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16091cs.LG

面向智能体What-If推理的基础模型蒸馏:混合LLM+SLM架构中的成本、延迟与治理

Distilling Foundation Models for Agentic What-If Reasoning:Cost, Latency, and Governance in a Hybrid LLM+SLM Architecture

Sourish Dey, Aditya Kumar

首次发表
浏览论文内容

中文总结 AI 辅助

针对表格基础模型推理延迟高的问题,通过将TabPFN蒸馏为紧凑前馈模型,在保持高准确率的同时实现数千倍参数压缩,并验证软标签蒸馏的有效性。

中文摘要 AI 辅助

表格基础模型通过上下文学习提供了强大的零训练预测性能,但其高推理延迟使其在交互式智能体循环中作为热路径决策后端不切实际。我们在UCI Adult和五个OpenML基准上的业务决策模拟中,将TabPFN教师模型蒸馏为紧凑的前馈学生模型:分类头将参数从53.2M压缩至8,546(6,220倍);部署的双头贷款流水线将参数从111.4M压缩至17,059(6,532倍)。学生模型保持了95.4-100.5%的准确率和96.8-100.0%的AUC,其中在credit-g上准确率保留最低为95.4%;一个alpha=0的硬标签对照实验表明,教师模型的软目标提供了2.1-7.0个AUC点的增益。

英文摘要

Tabular foundation models deliver strong zero-training predictive performance via in-context learning, but their high inference latency makes them impractical as hot-path decision backends in interactive agentic loops. We distill a TabPFN teacher into a compact feed-forward student across a business-decision simulation on UCI Adult and five OpenML benchmarks: the classification head compresses 53.2M parameters to 8,546 (6,220x); the deployed two-head loan pipeline compresses 111.4M parameters to 17,059 (6,532x). The student retains 95.4-100.5% accuracy and 96.8-100.0% AUC, with the lowest accuracy retention on credit-g at 95.4%; an alpha = 0 hard-label control shows that the teacher's soft targets provide a 2.1-7.0 AUC point gain.

发表机构

  • SumUp
  • Johannes Gutenberg-Universität Mainz(美因茨约翰内斯·古腾堡大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑