发表机构
Polytechnique Montreal(蒙特利尔理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究比较了函数令牌与提示中模式两种函数表示方法,在四个小语言模型上评估车辆函数调用,发现模式表示能泛化到未见函数并更可靠拒绝,其作用超过模型规模。
AI 中文摘要
车载助手必须在严格的内存和延迟约束下将自然语言请求转换为准确的车辆函数调用,这使得小语言模型(SLM)适合设备端部署。对于此类模型,一个关键的设计选择是如何呈现可用的函数表面。两种方法是:用专用的函数令牌(FT)表示每个函数,或直接在提示中提供函数模式。FT支持紧凑推理,但仅限于训练期间学到的函数,而提示中模式(SIP)可以泛化到未见过的函数,代价是更长的提示和更高的推理开销。我们引入了一个基准,包含9,822个单轮示例,涵盖源自Android Automotive的79个车辆函数,包括留出函数和需要拒绝的请求。我们在四个参数规模从270M到1.7B的SLM上,在匹配微调下比较了这两种方法。在训练期间见过的函数上,扩展带来的收益有限:270M模型可以匹配1.7B模型,而最强的整体性能出现在0.6B。在留出函数上,FT按构造达到零准确率,而SIP能泛化并随规模显著提高。在范围外请求上,FT可能调用它被训练过但不可用的函数,而SIP更可靠地基于提供的函数进行拒绝。这种灵活性带来更高的内存使用和延迟。我们的理论分析解释了SIP如何实现泛化,以及为什么更长的模式上下文会增加推理成本。总体而言,函数表面表示,而非仅模型规模,决定了基于SLM的车辆函数调用的能力和失败模式。
英文摘要
In-vehicle assistants must translate natural-language requests into accurate vehicle function calls under strict memory and latency constraints, making small language models (SLMs) attractive for on-device deployment. For such models, a key design choice is how the available function surface is presented. Two approaches are to represent each function with a dedicated Functional Token (FT) or provide function schemas directly in the prompt. FTs enable compact inference but are restricted to functions learned during training, whereas Schema-in-Prompt (SIP) can generalize to unseen functions at the cost of longer prompts and higher inference overhead. We introduce a benchmark of 9,822 single-turn examples spanning 79 vehicle functions derived from Android Automotive, including held-out functions and requests requiring refusal. We compare both approaches under matched fine-tuning across four SLMs from 270M to 1.7B parameters. On functions seen during training, scaling provides limited benefit: the 270M model can match the 1.7B model, and the highest Seen accuracy in our main grid occurs at 0.6B. On held-out functions, FT achieves zero accuracy by construction, whereas SIP generalizes and improves substantially with scale. On out-of-scope requests, FT can invoke an unavailable function it was trained to emit, while SIP more reliably refuses based on the functions offered. This flexibility comes with a higher inference cost, chiefly the latency of processing the schema prompt. Our theoretical analysis formalizes why only SIP can predict functions withheld as training targets and why longer schema contexts increase inference cost. Overall, function-surface representation, rather than model scale alone, determines the capabilities and failure modes of SLM-based vehicle function calling.
CommentsCamera-ready version. Accepted at the SLM-Agents Workshop at NeurIPS 2026 (non-archival)