arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TRACE:用于企业语言模型中知识保留参数化工具检索的业务规则基础推理课程

TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs

Sai Shruthi Sistla, Ashutosh Hathidara, Christopher Toukmaji, Mayank Shrivastava, Karthikeyan Asokkumar

arXiv 2607.22639首次发表:更新:

发表机构

SAP Labs(思爱普实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对企业语言模型中参数化工具检索问题,提出TRACE两阶段课程。第一阶段用多格式记忆SFT播种工具知识,第二阶段利用特定数据源训练模型生成思维痕迹,实现知识保留与单束贪婪解码,提升工具理解和检索召回率。

AI 中文摘要

参数检索使语言模型能够通过为每个 API 分配唯一虚拟令牌并通过约束束搜索训练模型来生成令牌来隐式检索工具。Toolsense表明这种方式有两个关键缺点:在训练过程中会破坏参数工具知识,并且其束搜索解码对于实时部署来说太慢。我们引入了TRACE(通过增强思维链和企业规则进行工具检索),这是一个两阶段课程来解决这种脱节问题。第一阶段重用来自ToolSense的多格式记忆SFT,用LoRA来播种工具知识。第二阶段是核心贡献:模型在生成工具令牌的JSON列表之前被训练以发出思维痕迹,使用两个数据源——来自ToolSense的RRB对和针对领域专家策划的目标业务规则合成的查询,两者都用推理痕迹增强。这种训练目标在生产延迟时实现单束贪婪解码的同时保留了第一阶段的MCQ和QA探测准确性。在两个企业产品线的8300多个工具的组合企业目录上进行评估,第二阶段的TRACE训练不仅保留而且提高了工具理解:MCQ准确性比第一阶段提高了3.2个百分点,QA探测提高了9个百分点。在检索方面,TRACE在领域A上实现了约86%的召回率,在领域B上实现了约60%的召回率,而嵌入基线性能分别约为27%和52%,两者都采用单束贪婪解码,使其能够在生产延迟时直接部署。

英文摘要

Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and training the model to generate it via constrained beam search. Toolsense shows that this regime has two critical drawbacks: it destroys parametric tool knowledge during training, and its beam-search decoding is too slow for real-time deployment. We introduce TRACE (Tool Retrieval via Augmented Chain-of-thought and Enterprise rules), a two-stage curriculum that resolves this dissociation. Stage 1 reuses the multi-format memorization SFT from ToolSense to seed tool knowledge with LoRA. Stage 2 is our core contribution: the model is trained to emit a thinking trace before producing a JSON list of tool tokens, using two data sources -- RRB pairs from ToolSense and queries synthesized to target business rules curated by domain experts -- both augmented with reasoning traces. This training objective preserves Stage 1 MCQ and QA probing accuracy while enabling single-beam greedy decoding at production latency. Evaluated on a combined enterprise catalog of 8,300+ tools across two enterprise product lines, TRACE training for Stage 2 not only preserves but improves tool understanding: MCQ accuracy gains +3.2 pp and QA probing gains +9 pp over Stage 1. On retrieval, TRACE achieves ~86% recall on Domain A and ~60% on Domain B -- compared to embedding baseline performance of ~27% & ~52% -- both with single-beam greedy decoding, making it directly deployable at production latency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑