arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自信地出错,悄无声息:对已部署的端侧语言模型不可检测故障的审计

Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model

Shashwat Pandey, Satwik Pandey, Suresh Raghu

arXiv 2608.23663首次发表:更新:

发表机构

University of California, Santa Cruz(加州大学圣克鲁兹分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究对已部署的端侧语言模型开展可靠性审计,发现其存在任务非对称校准误差等不可检测故障,提出模型无关审计协议等基础设施,黑盒一致性包装器可有效恢复其可靠性。

AI 中文摘要

对齐已部署的语言模型需要知晓其输出何时可信,但如今端侧模型被发送至数亿台设备,且无服务器端审核,开发者实际可部署的配置极少被独立审计。我们对开发者可访问的端侧基础模型开展了可复现的可靠性审计,核心问题为:用户或资源受限的开发者能否判断模型何时出错?通过对其校准进行红队测试、对错误前提问题的自信编造,以及对良性提示的过度拒绝,我们发现了一种任务非对称校准误差:其安全护栏在不同任务上向相反方向失效(对69%的错误前提进行编造,同时拒绝18%的完全良性输入),且模型自报告的置信度饱和且无区分性(AUROC为0.47;ECE为70,在可比小型模型中最差)。关键的是,自信正确与自信错误的输出在表面上无法区分:基于15项用户可见特征的分类器仅以0.55的AUROC区分二者(已证实等价),推理时无监督信号。无廉价的单轮生成信号可标记这些故障(AUROC≤0.68),而无需模型访问的黑盒一致性包装器可恢复可靠性(自信编造从75%降至3%;选择性准确率从43%升至83%),且可调整成本。我们贡献了一种模型无关的审计协议、一种表面不可区分性测试,并发布了代码与冻结评估项,作为审计已部署模型的可复用基础设施。

英文摘要

Aligning deployed language models requires knowing when their outputs can be trusted, yet on-device models now ship to hundreds of millions of devices with no server-side moderation, and the configuration developers can actually deploy is rarely audited independently. We present a reproducible reliability audit of the developer-accessible on-device foundation model, framed as an oversight question: can a user or a resource-constrained developer tell when the model is wrong? Red-teaming it on calibration, confident confabulation on false-premise questions, and over-refusal of benign prompts, we find a \emph{task-asymmetric miscalibration}: its guardrails fail in opposite directions across tasks (confabulating on 69\% of false premises while refusing 18\% of entirely benign inputs), atop a self-reported confidence that is saturated and non-discriminative (AUROC 0.47; ECE 70, worst among comparable small models). Crucially, confident-correct and confident-wrong outputs are \emph{surface-indistinguishable}: a classifier over 15 user-visible features separates them at AUROC only 0.55 (equivalence-confirmed), leaving no signal for oversight at inference time. No cheap single-generation signal flags these failures ($\le$0.68 AUROC), whereas a black-box consistency wrapper requiring no model access recovers reliability (confident confabulation 75\%$\to$3\%; selective accuracy 43\%$\to$83\%) at a tunable cost. We contribute a model-agnostic audit protocol, a surface-indistinguishability test, and released code and frozen evaluation items as reusable infrastructure for auditing deployed models.

Comments8 pages, 5 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑