发表机构
CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AgentProv是首个基于动作的智能体LLM API身份审计方法,通过工具调用分布指纹结合MMD置换测试,实现100%替换模型检测且误报率低,可有效应对现有审计方法的缺陷。
AI 中文摘要
商业大语言模型(LLM)API会宣传特定的基础模型,但实际提供的骨干模型可能被悄悄替换、量化或包装,例如为了节省部署成本。所有现有的审计方法都从文本输出通道判断骨干模型的身份,这对智能体API来说存在结构脆弱性,因为现代服务栈(OpenAI、Anthropic、Gemini、Cloudflare Workers AI、LangGraph)会丢弃文本,仅在模型调用工具时暴露结构化操作,且提供商注入的系统提示可能会扭曲文本分布,导致文本通道测试错误地指责诚实的提供商替换了声称的模型。我们观察到,近期的智能体后训练已将工具使用直接内化到模型权重中,这开辟了一个服务栈仍会暴露且对部署环境基本不变的新审计通道。我们提出Agentic Provenance(AgentProv),这是首个基于动作的智能体LLM API身份审计方法:AgentProv通过分类工具调用分布对部署模型进行指纹识别,并通过MMD置换测试判断身份。AgentProv在630个评估的检查点对上对所有替换模型的检测准确率达100%,同时在系统提示注入下的误报率低于7%(相比之下,MET的误报率为67%,RUT为53%)。在第三方API端点上,AgentProv与MET的分歧与检测提供商注入系统提示的独立令牌计数侧通道结果一致。
英文摘要
Commercial LLM APIs advertise a specific foundation model, but the served backbone may be silently substituted, quantized, or wrapped, for example to save deployment costs. All existing audits decide backbone identity from the text-output channel, which is structurally fragile for agentic APIs because modern serving stacks (OpenAI, Anthropic, Gemini, Cloudflare Workers AI, LangGraph) discard text and expose only structured actions when the model calls a tool, and provider-injected system prompts can distort text distributions enough that text-channel tests falsely accuse honest providers of substituting the claimed model. We observe that recent agentic post-training internalizes tool-use directly into the weights, opening a new audit channel that the serving stack still exposes and that is largely invariant to deployment context. We introduce Agentic Provenance (AgentProv), the first action-based identity audit for agentic LLM APIs: AgentProv fingerprints a deployed model through its categorical tool-call distribution and decides identity via an MMD permutation test. AgentProv catches every substituted model (100% on 630 evaluated checkpoint pairs), while holding the false-positive rate under system-prompt injection at 7% (vs. 67% for MET and 53% for RUT). On third-party API endpoints, AgentProv's disagreements with MET are consistent with an independent token-count side-channel that detects provider-injected system prompts.
Comments15 pages, 4 figures. Accepted to EMNLP 2026