哪个模型真正在为你服务?IRIS:大语言模型网关中模型替换和路由稀释的预算黑盒审计
Which Model Is Actually Serving You? IRIS: Budgeted Black-Box Auditing of Model Substitution and Routing Dilution in LLM Gateways
浏览论文内容
中文总结 AI 辅助
研究大语言模型网关中模型替换和路由稀释问题,提出IRIS审计方法,仅需返回文本,能结合多种检测与估计,通过实验验证其在不同场景下的有效性,相比其他审计方法有优势,还进行了多方面拓展实验。
中文摘要 AI 辅助
商业大语言模型网关介导对托管模型的访问,但提供服务的后端可能与宣传的不符:它可能在每个请求上替换更便宜的模型,或者仅将一小部分(ε)请求路由到该模型。先前的黑盒审计通常需要特权信号(对数概率、令牌排名或参考样本)或特定目标的探针,提前固定查询预算,并返回是/否的裁决。我们提出了IRIS审计,它只需要返回的文本:它要求端点生成随机数或字符串,对后端进行指纹识别,并且是第一个在仅文本的审计中,结合检测全流替换和分数稀释、确定提供服务的后端、估计路由分数(ε)以及自行确定的查询预算。一个低成本的试点拟合指数查询误差衰减,并在发出任何可疑查询之前冻结该预算。在Qwen3家族内部阶梯上,IRIS以0.99的曲线下面积(AUROC)验证后端,并随着查询的积累锐化归属;在商业OpenRouter库中,它在边际合格对中捕获ε = 0.3的稀释,平均功率为0.85(误报率为0.017),并将已注册稀释剂的ε恢复到0.04以内;实时跨提供商审计通过真正的量化和内核偏差标记15个同模型提供商对中的14个,在第三方MET跟踪中得到证实。与可比的黑盒审计相比,IRIS在共享任务上的检测匹配或更好,自适应分配将匹配预算目标命中率从73%提高到87%。进一步的实验涵盖对抗性网关、旋钮可识别性、未见过的稀释剂和误报控制。
英文摘要
Commercial LLM gateways mediate access to hosted models, but the served backend may not match the advertised one: it may substitute a cheaper model on every request or route only a fraction $ε$ of requests to it. Prior black-box auditors often need a privileged signal (log-probabilities, token ranks, or reference samples) or a target-specific probe, fix the query budget in advance, and return a yes/no verdict. We present $\mathrm{IRIS}$, an audit that needs only the returned text: it asks endpoints to generate random numbers or strings, fingerprints the backend, and is the first to combine, in one text-only audit, detection of whole-stream substitution and fractional dilution, attribution of the served backend, routing-fraction ($ε$) estimation, and a query budget it sizes itself. A cheap pilot fits the exponential query-error decay and freezes that budget before any suspect query is issued. On an intra-family Qwen3 ladder $\mathrm{IRIS}$ verifies the backend at $0.99$ AUROC and sharpens attribution as queries accumulate; across a commercial OpenRouter library it catches $ε{=}0.3$ dilution on margin-qualified pairs at $0.85$ mean power ($0.017$ false-positive rate) and recovers $ε$ to within $0.04$ for enrolled diluents; and a live cross-provider audit flags $14$ of $15$ same-model provider pairs by genuine quantization and kernel deviations, corroborated on third-party MET traces. Against comparable black-box auditors, $\mathrm{IRIS}$ matches or beats detection on shared tasks, and adaptive allocation lifts the matched-budget target-hit rate from $73$% to $87$%. Further experiments cover adversarial gateways, knob identifiability, unseen diluents, and false-positive control.