arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

保真度≠安全性:轻度压缩的大语言模型通过所有无数据质量防护,但在智能体执行中生成指令中不存在的流程步骤

Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution

I. Kennedy, T. Kennedy

arXiv 2607.28196首次发表:更新:

AI 中文总结

该研究发现轻度压缩LLM虽通过困惑度、MMLU等无数据质量防护,但在智能体执行SOP时会生成指令外流程步骤,提出基于压缩误差双轴统计量的无数据筛查方法以保障智能体安全。

AI 中文摘要

从业者通常会接受满足一系列低成本数据质量防护的压缩语言模型,这些防护包括:困惑度与原模型的差距在可接受范围内、下游任务准确率(如MMLU)处于置信区间内,以及通过随机探测输入对比压缩模型与原模型内部表征的无数据输出保真度信号。但这套防护存在盲点:在三类模型上,轻度压缩模型通过所有防护后,作为智能体执行标准操作流程(SOP)时,会生成指令中从未包含的流程步骤。该效应与压缩操作相关:相干低秩(SVD)截断会引发此问题,而匹配相同困惑度的幅值剪枝则不会。一项分离实验揭示了原因:通过配对输出保真度测试(置信区间内)的相同压缩权重,在针对生成步骤的金丝雀测试中未通过置信区间检验。核心影响因素是压缩误差的相干性乘以其比率,而非损伤程度。无数据保真度探测本质是保真度 oracle,无法识别该维度。我们通过预注册、足够功效的金丝雀测试,在三类模型上用配对置信区间表征该盲点与分离效应,且该操作相关性在三类模型上均复现,困惑度防护的规避发生在模型存在防护内低秩余量的场景。我们随后提出一种无数据筛查方法:基于压缩误差的双轴统计量(相干分数与误差率),通过跨模型架构的固定阈值标记失效构建,契合相干×比率机制。困惑度、MMLU与保真度接受度无法证明智能体安全性,需在智能体部署前筛查轻度压缩的低秩构建。

英文摘要

Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and data-free output-fidelity signals that compare the compressed and original network's internal representations under random probe inputs. This stack has a blind spot. Across three model families, gently-compressed models clear every guard and then invent procedure steps that were never in the instructions when they run a standard operating procedure (SOP) as an agent. The effect is operator-specific: coherent low-rank (SVD) truncation induces it, and magnitude pruning matched to the same perplexity does not. One dissociation isolates the cause. The same compressed weights that CI-win a paired output-fidelity test CI-fail the invented-step canary. The governing axis is the coherence of the compression error times its rate; the magnitude of the damage does not predict it. The data-free fidelity probe is a fidelity oracle by construction, so it cannot see this axis. We characterize the blindspot and dissociation with paired confidence intervals on a pre-registered, powered canary across three architectures. Operator-specificity replicates on all three, and the perplexity-guard evasion appears where the model admits in-guard low-rank headroom. We then give a data-free screen: a two-axis statistic of the compression error (coherent-fraction and error-rate) that flags the failing builds with fixed thresholds across architectures and matches the coherence-times-rate mechanism. Perplexity, MMLU, and fidelity acceptance do not certify agent safety. Screen gently-compressed low-rank builds before agentic deployment

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑