arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

虚假底线:LLM安全路由评估在分布偏移下失效

False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift

Amit Singh Bhatti, Vishal Vaddina

arXiv 2610.01535首次发表:更新:

发表机构

Quantiphi Analytics(Quantiphi分析公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示LLM安全路由评估在分布偏移下因比较器选择偏差而失效,量化其危害并证明路由收益甚微,提出应在偏移下评估且防御应基于危害评分。

AI 中文摘要

安全路由器将每个请求发送给多个模型之一,并与最佳单一模型进行比较评判。一个主要的路由基准在评估数据上选择该比较器。在基准自身的设定下,这并无危害,但在分布偏移下则不然。在HELM Safety上,随机划分下的选择成本为0.003-0.030的危害,而保留类别下为0.045-0.113,与归因于路由的全部缺陷相当,且其方向在任一已发布的评判器下均成立。在AgentDojo上,当套件被保留时,该成本上升七到九倍。在预先按规则选定的七个安全语料库中,三个满足注册区间检验,四个击败了后续的排列零假设,四个区间未命中中有三个是某些模型观测到零危害的语料库。先前工作证明了该偏差的方向。我们量化了其对危害和准确性的影响,表明在我们测量的保留划分下影响更大,并通过乐观性加上依赖偏移的遗憾来界定。在偏移下诚实评分,路由在这些基准上收益甚微。在大多数池单元中,嵌套路由器服务于诚实基线的模型,而在几乎饱和的AgentDojo语料库上,一个完美的预调度路由器至多值两个点的危害。我们还发现模型对晚期注入的明确识别是可操纵的。在保留重跑中,知道面对哪个模型的攻击者将GPT-5.4的评判识别降低19.6个点,并由独立标签确认。在离线反事实组合到控制器中时,同样的攻击根据回退模型提高或降低估计危害。安全路由应在偏移下评估,针对无测试标签选择的基线,基于识别的防御应在攻击者选择模型所见内容的危害上评分。

英文摘要

Safety routers send each request to one of several models and are judged against the best single model. A major routing benchmark picks that comparator on the evaluation data. In the benchmark's own setting this is harmless, but under distribution shift it is not. On HELM Safety the selection cost is 0.003-0.030 of harm under random splits and 0.045-0.113 under held-out categories, comparable to the whole deficit attributed to routing, with its direction holding under either published judge alone. It rises seven- to ninefold on AgentDojo when suites are held out. Across seven safety corpora chosen by rules fixed in advance, three meet a registered interval test and four beat a later permutation null, and three of the four interval misses are corpora where some models have zero observed harm. Prior work proves the direction of this bias. We size it on harm and accuracy, show that it is larger under the held-out splits we measure, and bound it by optimism plus a shift-dependent regret. Scored honestly under shift, routing buys little on these benchmarks. In most pool cells the nested router serves the honest baseline's model, and on the nearly saturated AgentDojo corpus a perfect pre-dispatch router is worth at most two points of harm. We also find a model's expressed recognition of a late injection steerable. On held-out reruns an attacker who knows which model it faces lowers GPT-5.4's judged recognition by 19.6 points, confirmed by an independent label. In an offline counterfactual composition into a controller, the same attack raises or lowers estimated harm depending on the fallback model. Safety routing should be evaluated under shift, against a baseline chosen without the test labels, and recognition-based defences should be scored on harm against an attacker who chooses what the model sees.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑