arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23886cs.CL

this-that-model-1.0:一个30毫秒内决策、成本为百万分之一美分的类型化决策模型

this-that-model-1.0: A typed decision model that decides in 30 ms, for a millionth of a cent

  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

Zehua Cheng, Wei Dai, Jiahao Sun

AI总结:

本文提出20亿参数的this-that-model-1.0类型化决策模型,直接从隐藏状态读取答案,30.9毫秒内决策且零token生成,成本极低,在68个决策问题上准确率0.941,优于托管服务,但多步算术能力有限。

AI中文摘要:

软件每年将更多的分支决策委托给模型:工单进入哪个队列、命令是否安全执行、索赔是否无需人工处理。程序需要的不是散文,而是n个声明选项中的一个,以及一个可以设置阈值的数字。如今,这需要往返调用前沿模型——数百毫秒、按token计费,以及一个解析器——而问题通常只是三个子句的合取。this-that-model-1.0是一个20亿参数的类型化决策模型。其答案直接从指定位置的隐藏状态读取,并限制在调用者声明的选项集内,因此不生成文本,不会产生格式错误,并且请求中的每个问题都在同一次前向传播中得到回答。它在单个笔记本电脑GPU上30.9毫秒内做出决策,且生成零个输出token,而前沿API调用需要8758毫秒,且能很好回答这些问题的托管系统每个问题先思考,生成21到212个token,每个token都计费。它在单个消费级GPU上每秒维持32次决策,且状态从不离开机器。在第三方记录的68个决策问题队列上,基于他们的输入和措辞,其得分为0.941,Brier分数为0.042,而托管服务Jev在相同项目上的得分分别为0.765和0.133。我们内部42个任务族套件的一次运行需要32秒,电费为0.000217美元;我们测得的最高准确率托管模型需要155.2分钟和10.636美元。我们还报告了其不足之处。在多步算术上,由于单次前向传播无法携带中间结果,其得分为0.560,而它们的得分为0.98至1.00;针对性的第二轮训练提高了其编写时针对的五个任务族,但未迁移到其他13个任务族。该模型在此https URL开源。

英文摘要:

Software delegates more of its branches to models every year: which queue a ticket enters, whether a command is safe to run, whether a claim clears without a person. What the program needs back is not prose. It is one of n declared options and a number it can threshold. Today that costs a round trip to a frontier model -- hundreds of milliseconds, a per-token bill, and a parser -- for a question that is usually a conjunction of three clauses. this-that-model-1.0 is a 2B-parameter typed decision model. Its answer is read directly from the hidden state at a designated position and restricted to the option set the caller declared, so no text is generated, nothing can be malformed, and every question in a request is answered in the same forward pass. It decides in 30.9 ms on one laptop GPU and generates zero output tokens doing it, where a frontier API call costs 8758 ms and the hosted systems that answer these questions well spend between 21 and 212 generated tokens per question thinking first, billed for every one. It sustains 32 decisions per second on one consumer GPU and never lets the state leave the machine. On a third party's recorded cohort of 68 decision questions, on their inputs and their wording, it scores 0.941 with a Brier score of 0.042, against 0.765 and 0.133 for the hosted service Jev on the same items. One pass of our 42-family internal suite takes 32 seconds and 0.000217 USD of electricity; the most accurate hosted model we measured needs 155.2 minutes and 10.636 USD. We also report where it loses. On multi-step arithmetic, which a single forward pass cannot carry intermediate results through, it scores 0.560 against their 0.98 to 1.00, and a targeted second training round improved the five task families it was written for and transferred to none of the other 13. The model is open-sourced in https://huggingface.co/flock-io/this-that-model-1.0

↑