HARISSA:推理时自检实现高效安全的本地语言模型部署
HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment
浏览论文内容
中文总结 AI 辅助
HARISSA利用模型隐藏状态在推理时进行自检,通过级联策略在效率和安全性间权衡,实现本地部署下低延迟高准确率的智能决策。
中文摘要 AI 辅助
在本地运行语言模型在隐私、延迟和成本方面具有优势,但本地硬件只能容纳小模型,其能力不如前沿模型。对于困难查询,通常的补救措施是将其升级到云端模型,但这会放弃本地运行带来的隐私和成本优势。相反,保持本地的部署在面临困难查询时需要做出两个决策。首先,它可以在查询上花费更多计算,例如在回答前进行推理,这会以提高延迟为代价提高准确性,因此必须决定哪些查询值得额外计算(效率)。其次,某些查询超出了本地模型的能力,而提供错误答案比将查询转交给人类介入更糟糕,因此必须决定哪些答案可以安全交付(安全性)。我们表明,这两个决策都可以从模型自身的隐藏状态中得出。预填充状态(在任何令牌生成之前计算)预测模型是否会正确回答,而答案状态(在生成的答案结束时)预测该答案是否正确。HARISSA对模型进行微调,使这两种状态都能预测正确性,然后通过一个级联策略做出这两个决策,该策略按从最便宜到最昂贵的回答方式依次进行,跳过预填充状态预测会失败的方式,并在其停止时的答案被预测为错误时推迟查询。在运行单个模型的设备上,HARISSA的准确率与思维链相差一个百分点以内,而延迟降低2.7倍。在持有同一模型四种规模的服务器上,HARISSA在相同延迟下比FrugalGPT和Self-REF级联更准确,并且在相同的延迟率下,答案状态在六个任务和设置组合中的五个中比标准置信度信号留下更少的错误答案。
英文摘要
Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally. A deployment that stays local faces two decisions for hard queries instead. First, it can spend more computation on a query, e.g., reasoning before answering, which raises accuracy at a cost in latency, so it must decide which queries are worth the extra computation (efficiency). Second, some queries are beyond the local model, and delivering a wrong answer is worse than deferring the query to a human in the loop, so it must decide which answers are safe to deliver (safety). We show that both decisions can be made from the model's own hidden states. The prefill state, computed before any token is generated, predicts whether the model will answer correctly, and the answer state, at the end of the generated answer, predicts whether that answer is correct. HARISSA fine-tunes the model so that both states predict correctness, then makes both decisions with one policy that cascades through the ways of answering from cheapest to most expensive, skipping a way the prefill state predicts will fail and deferring the query when the answer it stops with is predicted wrong. On a device running a single model, HARISSA is within one accuracy point of chain-of-thought at 2.7 times lower latency. On a server holding four sizes of one model, HARISSA is more accurate than the FrugalGPT and Self-REF cascades at the same latency, and at the same deferral rate the answer state leaves fewer wrong answers than the standard confidence signals in five of six task and setting pairs.
发表机构
- University of Michigan(密歇根大学)
- SNU-LG AI Research Center(首尔大学-LG人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。