arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

所有权的形态:通过语义结构验证大语言模型(LLM)的出处

The Shape of Ownership: Verifying LLM Provenance through Semantic Structures

Zhongrui Sun, Jiahao Chen, Oubo Ma, Yuwen Pu, Zhou Feng, Haibo Hu, Shouling Ji

arXiv 2609.02553首次发表:更新:

AI 中文总结

该研究针对LLM所有权验证难题,提出PROSE方法,通过语义结构实现黑盒场景下的高鲁棒性、高隐蔽性的模型出处验证,在多类模型上实现100%检测率且无假阳性。

AI 中文摘要

随着大语言模型(LLM)被越来越多地重新分发、适配和部署在不透明的API背后,已无法通过检查模型内部结构或部署记录来可靠地确定模型所有权,这催生了对可通过黑盒交互观测的行为特征的需求。然而,大多数现有的黑盒指纹通过固定的查询-键关联来实例化所有权信号,将模型身份简化为与普通行为脱节的稀疏记忆关联,既限制了对下游修改(如微调或量化)的鲁棒性,也降低了隐蔽性。更强的指纹应是分布式的、自然触发的,并在更高的语义层面上表达。为此,我们提出PROSE(Provenance through Relational Organization of Semantic Expression,即通过语义表达的关系组织确定出处),用目标语义领域替代固定查询集,用语义结构(作为领域条件响应行为内化)替代脆弱的响应键。具体而言,该指纹编码在模型对其领域内结论的语义组织方式中,而非特定令牌或规定输出中。PROSE构建了一个领域特定语义模板的私有库,通过对结构验证的干净响应进行混合微调将其内化,并通过在保留的自然查询响应中检测指定结构来验证所有权。在多个模型架构、规模和目标领域上进行的广泛实验表明,PROSE在未修改的模型上实现了100%的指纹检测率,未观察到误报,同时保留了模型效用,并在下游修改和输出变换下保持强可检测性。

英文摘要

As large language models (LLMs) are increasingly redistributed, adapted, and served behind opaque APIs, model ownership can no longer be established reliably by inspecting model internals or deployment records. This creates a need for behavioral signatures that remain observable through black-box interaction. Yet most existing black-box fingerprints instantiate ownership signals through fixed query-key associations, reducing model identity to sparse memorized associations detached from ordinary behavior and limiting both robustness and stealth (e.g., fine-tuning or quantization) and stealthiness. A stronger fingerprint should instead be distributed, naturally elicited, and expressed at a higher semantic level. To this end, we introduce PROSE (Provenance through Relational Organization of Semantic Expression), replacing fixed query sets with a target semantical domain and brittle response keys with semantic structures internalized as domain-conditioned response behavior. Specifically, the fingerprint is encoded in how the model semantically organizes its in-domain conclusions, rather than in particular tokens or prescribed outputs. PROSE constructs a private bank of domain-specific semantic templates, internalizes them through mixed fine-tuning on structurally verified and clean responses, and verifies ownership by detecting the designated structures in responses to held-out natural queries. Extensive experiments across multiple model architectures, scales, and target domains show that PROSE achieves a 100% fingerprint detection rate on unmodified models with no observed false positives, preserves model utility, and retains strong detectability under downstream modifications and output transformations.

Comments20 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑