arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.31142cs.SEcs.AIcs.CR

匿名AI模型审计:一种用于黑盒身份验证的四阶段协议

Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification

Yisen Xi

首次发表
浏览论文内容

中文总结 AI 辅助

针对匿名AI模型黑盒身份验证的需求,提出四阶段取证审计协议,经测试可验证模型声明一致性,部分案例能准确推断模型身份。

中文摘要 AI 辅助

2025至2026年的AI市场出现了一波秘密发布浪潮:前沿模型以代号形式在开发者平台上匿名推出。对于用户而言,模型身份决定了数据处理条款、供应链风险和能力预期。目前尚不存在针对匿名模型的黑盒身份验证的验证方法:从业者的检查清单缺乏准确性证据,而模型的自我识别从设计上就不可信。我们提出了一种针对API服务模型的四阶段取证审计协议。第0阶段通过存档的平台快照(互联网档案馆)重建发布时的配置,揭示预览版本与生产版本之间的差异。第1阶段针对平台目录对配置(上下文、输出上限、推理能力、模态)进行指纹识别。第2阶段使用跨长度差分测试分词器身份,该差分可拒绝短提示碰撞。第3阶段通过行为探测进行佐证。我们在10个已知身份的发布版本上测试了声明一致性(7个完全匹配,2个精度差异,1个部分匹配,0个反向匹配),未针对匿名性进行端到端识别。在一个旗舰案例中对识别进行了前瞻性验证,该案例在2026年8月23日的分析指向了GLM-5.3版本系列,其官方发布确认了这些系列和版本系列的推断(部署变体未预先声明;Flash在发布后保持一致),以及三个仅第0阶段的案例,其中协议产生了分级假设或弃权(不执行)而非猜测。补充材料提供了仅使用标准库的实现。

英文摘要

The 2025--2026 AI market has seen a wave of stealth releases: frontier models launched anonymously on developer platforms under codenames. For their users, identity determines data-handling terms, supply-chain risk, and capability expectations. No validated methodology exists for black-box identity verification of anonymous models: practitioner checklists lack accuracy evidence, and self-identification is untrustworthy by design. We propose a four-stage forensic audit protocol for API-served models. Stage 0 reconstructs launch-time configuration from archived platform snapshots (Internet Archive), exposing preview--production drift. Stage 1 fingerprints configuration (context, output ceiling, reasoning, modality) against the platform catalog. Stage 2 tests tokenizer identity with a cross-length differential that rejects short-prompt collisions. Stage 3 corroborates with behavioral probes. We test declaration consistency on 10 known-identity releases (7 exact, 2 precision-differences, 1 partial, 0 counter-directional), not end-to-end identification under anonymity. Identification is validated prospectively on a flagship case whose 2026-08-23 analysis pointed to the GLM-5.3 version line and whose official reveal confirmed those family and version-line inferences (deployment variant was not pre-asserted; Flash was consistent post-reveal), and on three Stage-0-only cases where the protocol produced a graded hypothesis or declined rather than guessed. A standard-library-only implementation is provided as supplementary material.

发表机构

  • Independent Researcher, Beijing, China

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑