arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16391cs.CRcs.AI

Ventor-QTest:威胁模型驱动的供应商托管大语言模型API验证

Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs

  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

Xiangfan Wu, Zonghao Ying, Huiyu Wu, Xing Zheng, Huangsheng Cheng, Xiaorong Shi, Jing Guo

中文总结 AI 辅助

Ventor-QTest是一种无需目标API概率信息的复合黑盒审计方法,通过AFL和EFL指标审计供应商托管LLM API质量,可反映长视野任务的极端保真度损失,开源实现可获取。

中文摘要 AI 辅助

随着大语言模型的日益普及,部署开放权重模型的第三方供应商已成为生态系统的重要组成部分,因此对其推理API的质量进行审计是一个悬而未决的问题。我们将托管模型路由形式化为随机过程,并提出Ventor-QTest,这是一种复合黑盒审计,无需目标API提供概率信息。其重复请求组件将每个冻结的受限上下文发送到目标多次,从返回的文本计数中重建分类输出分布,并报告平均保真度损失(AFL),这是一种经零偏差校正的窗口内平均粗化KL统计量。其长序列组件使用独立运行,通过以运行为中心的参考惊奇统计量的经验上尾报告极端保真度损失(EFL)。在三个支持对数概率的路由条件下,AFL与基于对数概率的粗化KL比较器显示出强线性描述一致性。在七个路由快照中,20次运行的序列探测揭示了特定路由的EFL变化。AFL和EFL与GPQA-Diamond准确性几乎没有可检测到的路由级关联,相比之下,明显的EFL与Terminal-Bench通过率随任务暴露增加而下降相吻合,这种模式可能是因为长视野任务的正确性对极端保真度损失更敏感。这些结果促使联合报告AFL和EFL,特别是在审计长视野智能体任务时,开源实现可在该httpsURL获取。

英文摘要

As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part of the ecosystem. Auditing the quality of their inference APIs is therefore an open problem. We formalize hosted model routing as a stochastic process and propose \mbox{\textbf{Ventor-QTest}}, a composite black-box audit that requires no probability information from the target API. Its repeated-request component sends each frozen constrained context to the target multiple times, reconstructs a categorical output distribution from the returned text counts, and reports \emph{average fidelity loss} (AFL) as a null-bias-corrected, within-window mean coarsened-KL statistic. Its long-sequence component uses independent runs to report \emph{extreme fidelity loss} (EFL) through the empirical upper tail of a run-level reference-centered-surprisal statistic. Across three logprob-capable route conditions, AFL shows strong linear descriptive agreement with a logprob-derived coarsened-KL comparator. Across seven route snapshots, 20-run sequence probes reveal route-specific EFL variation. AFL and EFL have little detectable route-level association with GPQA-Diamond accuracy. In contrast, pronounced EFL coincides with a decline in Terminal-Bench pass rate as task exposure increases. This pattern may arise because correctness in long-horizon tasks is more sensitive to extreme fidelity loss. These results motivate reporting AFL and EFL jointly, particularly when auditing long-horizon agentic tasks. The open-source implementation is available at https://github.com/Tencent/AI-Infra-Guard/tree/main/services/api_checker/ventor_qtest.

↑