arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36115cs.LG

Koa-action:使用生成式大语言模型进行快速且一致的结构化决策

Koa-action: Fast and Consistent Structured Decision Making with Generative LLMs

  • Salesforce AI(赛富时人工智能研究院)
  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

机构由 AI 辅助整理,请以论文原文为准。

Shenghong Dai, Shiva Kumar Pentyala, Yingchi Liu, Shubham Mehrotra, Suman Banerjee, James Zhu, Bin Bi, Sitaram Asur, Phil Mui

AI总结:

针对LLM在延迟关键型分类中的不足,提出Koa-action框架,通过单令牌约束生成和微调实现快速原子决策,在保持高准确率的同时大幅降低延迟。

AI中文摘要:

行业应用通常要求低延迟分类,然而当前的大语言模型(LLM)方法仍不适合延迟关键型应用。现有的提示和约束解码会产生冗长的多令牌输出,需要昂贵的逐令牌生成,而基于编码器的模型虽然推理速度更快,但牺牲了任务灵活性。我们提出Koa-action,一个用于低延迟原子动作的框架——快速、单步决策,如分类、语义端点检测、布尔检查和评分——将其表述为具有单令牌输出的约束生成。通过引入原子标签令牌并应用监督微调,我们的方法将分类简化为确定性的单步解码问题。在标准基准测试中,Koa-action在保持一致低且稳定延迟的同时,达到了有竞争力的准确性。在一个生产意图路由基准上,Koa-action达到85.5%的准确率——与最强的前沿模型(Claude-4.8-Opus、Gemini-Pro-3.1)相当,并领先于GPT-5和Gemini-2.5-Pro——同时在相同服务条件下,回答时间约为半秒,比每个前沿模型快数倍(中位数快约7.5倍)。与专用的单令牌系统Jev/TypeSafe相比,Koa-action在准确性上具有竞争力,中位数速度更快,同时还处理了单标签文本系统无法处理的多模态输入和多标签输出。

英文摘要:

Industry applications often demand low-latency classification, yet current large language model (LLM) approaches remain poorly suited for latency-critical applications. Existing prompting and constrained decoding produce verbose, multi-token outputs that require expensive token-by-token generation, while encoder-based models achieve faster inference but sacrifice task flexibility. We propose Koa-action, a framework for low-latency atomic actions -- fast, single-step decisions such as classification, semantic endpointing, Boolean checks, and scoring -- formulated as constrained generation with single-token outputs. By introducing atomic label tokens and applying supervised fine-tuning, our method reduces classification to a deterministic one-step decoding problem. Across standard benchmarks, Koa-action delivers competitive accuracy with consistently low and stable latency. On a production intent-routing benchmark, Koa-action reaches 85.5% accuracy -- competitive with the strongest frontier models (Claude-4.8-Opus, Gemini-Pro-3.1) and ahead of GPT-5 and Gemini-2.5-Pro -- while answering in about half a second, several-fold faster than every frontier model (up to ~7.5x at the median) under identical serving conditions. Against the dedicated single-token system Jev/TypeSafe, Koa-action is competitive on accuracy and faster at the median, while also handling multimodal inputs and multi-label outputs that single-label text systems do not.

补充信息

↑