AI 中文总结
研究针对大语言模型易被滥用问题,提出统一指纹框架。通过LCF按规则构造代码混合指纹,再用LCFEdit结合多语言表示和跨语言对齐注入指纹,实现构造感知注入,确保更新稳定,能持续验证所有权且对模型效用影响小。
AI 中文摘要
大语言模型是昂贵的知识资产,易遭受未经授权的重新分发和商业滥用。注入指纹提供了一种实用的、黑盒可验证的所有权信号,但现有方法将指纹生命周期的构造和注入两个阶段解耦。现有指纹框架有两个局限性:自然语言指纹易意外激活,乱码指纹易被基于困惑度的检测过滤。此外,构造与注入解耦使注入阶段不知触发词的语言结构,错失针对性优化机会。我们提出一个统一指纹框架,联合优化两个阶段。首先,LCF通过语义密度替换规则和语法偏向混合来构造代码混合指纹,避免自然语言触发词的意外激活失败。其次,LCFEdit利用高资源多语言表示的零空间投影注入每个指纹,并通过跨语言对齐步骤增强,使权重更新朝向指纹语言的表示子空间。这种构造感知注入确保更新在语言上更稳定。广泛评估表明该方法能持续进行所有权验证,对效用影响可忽略不计。
英文摘要
Large language models (LLMs) are costly intellectual assets that remain exposed to unauthorized redistribution and commercial misuse. Injected fingerprints, i.e., trigger--target pairs embedded in model behavior, offer a practical, black-box-verifiable ownership signal, but existing methods decouple the two stages of the fingerprint life cycle: how a fingerprint is constructed and how it is injected. Existing fingerprinting frameworks suffer from two limitations. Natural-language fingerprints are prone to accidental activation, and garbled fingerprints are easily filtered by perplexity-based detection. Furthermore, decoupling construction from injection leaves the latter unaware of the trigger's linguistic structure, missing the opportunity for targeted optimization. We argue that fingerprint construction should drive injection, and present a unified fingerprinting framework that jointly optimizes both stages. First, LCF constructs code-mixing fingerprints by combining low-resource languages under a semantic-density substitution rule and grammar-biased mixing, yielding triggers whose perplexity sits far below garbled baselines while avoiding the accidental-activation failures of natural-language triggers. Second, LCFEdit injects each fingerprint with a null-space projection derived from high-resource multilingual representations that preserves knowledge, augmented by a cross-lingual alignment step that steers the weight update toward the fingerprint language's representation subspace. This construction-aware injection ensures that the update is linguistically informed and therefore more stable. Extensive evaluations on imperceptibility, detectability, and harmlessness demonstrate persistent ownership verification with negligible impact on utility.