参数高效微调(PEFT)适配何时会泄露结构?测量公共基础模型服务中的黑盒结构边界
When Do PEFT Adaptations Leak Structure? Measuring Black-Box Structural Bounds in Public-Base Model Services
浏览论文内容
中文总结 AI 辅助
本研究提出VectorHijack-SR方法,测量公共基础模型PEFT适配的结构泄露边界,发现PEFT服务会泄露结构与版本信息,且存在可见性-可利用性差距。
中文摘要 AI 辅助
服务越来越多地部署带有私有参数高效微调(PEFT)适配的公共基础模型,当审计员或对手可以在本地执行公共基础模型并观察受害者输出时,会产生差分信息泄露风险。我们提出VectorHijack-SR,一种测量方法,可将配对的受害者/基础残差转换为PEFT家族、层局部性和粗秩的校准结构边界,同时将元数据可见性与开放世界有效性和操作可利用性分离。我们的估计量将查询级别的幅度、排名、熵、间隔、长度、模板、局部性和光谱统计聚合为服务级表示。服务不相交分类器量化结构证据,交叉拟合分层弃权(不执行)评估器判断受害者是否位于校准的LoRA流形之外。在所有分类主干上,BERT/MNLI(8/12)、RoBERTa/MNLI(21/24)、DeBERTa-v3在MNLI(12/18)和AG News(15/18)上的家族泄露超过均匀随机概率。秩推断具有任务依赖性:BERT/MNLI和DeBERTa/AG News达到8/9,而DeBERTa/MNLI达到4/9,校正后与随机概率统计兼容。在十次随机种子的BERT开放集网格上,弃权(不执行)评估器的合并AUROC为0.804(95%置信区间[0.660, 0.927]),已知准确率为0.956,但在结构接近的DoRA和LoRA+head变体上仍有限。在五个保留的LoRA-r64服务上的精确版本链接达到AUC 0.940。实验揭示了可见性-可利用性差距:两阶段恢复未提供公平预算查询节省,后验选择的PEFT表现逊于蒸馏后转换的PEFT(0.356对比0.517),自由运行生成仍接近随机概率。这些结果表明,具有已知基础、丰富输出的PEFT服务可泄露可操作的结构和版本信息,而仅封闭集置信度无法确立通用适配器恢复能力。
英文摘要
Services increasingly deploy public foundation models with private parameter-efficient adaptations, creating a differential information leakage risk when auditors or adversaries can execute the public base model locally and observe victim outputs. We present VectorHijack-SR, a measurement methodology that converts paired victim/base residuals into calibrated structural bounds over PEFT family, layer locality, and coarse rank, while separating metadata visibility from open-world validity and operational exploitability. Our estimator aggregates query-level magnitude, ranking, entropy, margin, length, template, locality, and spectral statistics into service-level representations. A service-disjoint classifier quantifies structural evidence, and a cross-fitted hierarchical rejector evaluates whether a victim lies outside the calibrated LoRA manifold. Across classification backbones, family leakage exceeds uniform chance on BERT/MNLI (8/12), RoBERTa/MNLI (21/24), and DeBERTa-v3 on MNLI (12/18) and AG News (15/18). Rank inference is task dependent: BERT/MNLI and DeBERTa/AG News reach 8/9, whereas DeBERTa/MNLI achieves 4/9 and is statistically compatible with chance after correction. On a ten-seed BERT open-set grid, the rejector achieves pooled AUROC 0.804 (95% CI [0.660, 0.927]) and known accuracy 0.956, but remains limited on structurally close DoRA and LoRA+head variants. Exact-version linkage on five held-out LoRA-r64 services reaches AUC 0.940. Experiments reveal a visibility-exploitability gap: two-stage recovery provides no fair-budget query savings, posterior-selected PEFT underperforms distill-then-convert PEFT (0.356 vs. 0.517), and free-running generation remains near chance. These results show that known-base, rich-output PEFT services can leak actionable structural and version information, while closed-set confidence alone does not establish universal adapter recovery.