arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14883cs.CR

可逆性验证的云-本地大语言模型推理去标识化:一种具有分层保证的本地认证脱水-再水化循环(DR-SL)

Reversibility-Verified De-identification for Cloud-Local LLM Inference: A Locally Certified Dehydrate-Rehydrate Loop with Layered Assurance (DR-SL)

Wen Hu, Ya Yu, Xutong Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对云-本地LLM推理,提出DR-SL去标识化框架,通过双分支验证器迭代脱水并辅以硬线和人工审查,在保证安全的同时保持效用,显著降低泄漏并实现自动发布。

中文摘要 AI 辅助

云-本地大语言模型推理必须在利用云端级推理能力的同时,将敏感用户数据保留在设备端,然而现有的净化方法(占位符替换、差分隐私扰动和技能蒸馏)缺乏一种同时保证安全性和效用保持的发布决策。我们提出DR-SL(带自学习循环的脱水-再水化),将去标识化完整性形式化为两个可测量的条件:在Pufferfish语义下的去标识化充分性,以及通过QA探针进行的任务信息保持。一个完全本地的双分支验证器在具有保证终止的字典序门控下迭代脱水,并辅以确定性硬线、外部强攻击者重测和人工回退。我们证明了Fano型下界、一个Pufferfish见证以及一个速率-隐私可行性准则,并明确说明其适用范围:这些界证明泄漏,而非安全性,并且在我们运行点附近近乎空洞,因此发布安全性依赖于经验校准、硬线和人工审查。在一个最坏情况的全任务耦合基准上,该循环将泄漏从0.457降至0.304(p约等于0),发布链在出口处实现0.000字面泄漏(160个实例,两个强攻击者),系统恰好按照可行性准则预测的那样退化为认证和路由。在一个混合耦合基准上,相同的安全点在零测量泄漏下自动发布67.5%的实例,在相同发布规则下帕累托支配占位符和选择性LDP角点。两项人工研究锚定了语义效用度量(Spearman rho = 0.839)和注释金标准(类型级召回率至少0.987)。探索性自学习假设未得到支持,并如实报告。所有理论界均通过数值验证;代码、合成数据集、协议和人工研究包均已公开。

英文摘要

Cloud-local LLM inference must keep sensitive user data on-device while exploiting cloud-grade reasoning, yet existing sanitization approaches (placeholder substitution, differential-privacy perturbation, and skill distillation) lack a release decision that is simultaneously safe and utility-preserving. We propose DR-SL (Dehydrate-Rehydrate with Self-Learning loop), which formalizes de-identification completeness as two measurable conditions: de-identification sufficiency under Pufferfish semantics, and task-information preservation via QA probes. A fully local two-branch verifier iterates dehydration under a lexicographic gate with guaranteed termination, backed by a deterministic hard line, an external strong-attacker re-test, and human fallback. We prove Fano-type lower bounds, a Pufferfish witness, and a rate-privacy feasibility criterion, and state their scope plainly: the bounds certify leakage, never safety, and are near-vacuous at our operating point, so release safety rests on empirical calibration, the hard line, and human review. On a worst-case fully task-coupled benchmark the loop reduces leakage from 0.457 to 0.304 (p approx. 0) and the release chain delivers 0.000 literal leakage at egress (160 instances, two strong attackers), the system degrading to certification-and-routing exactly as the feasibility criterion predicts. On a mixed-coupling benchmark the same safe point releases 67.5% of instances automatically at zero measured leakage, Pareto-dominating placeholder and selective-LDP corners under an identical release rule. Two human studies anchor the semantic utility metric (Spearman rho = 0.839) and the annotation gold (type-level recall at least 0.987). The exploratory self-learning hypothesis was not supported and is reported as such. All theoretical bounds pass numerical verification; code, synthetic datasets, protocol, and human-study packages are public.

发表机构

  • Jiangsu Yunhefeng Intelligent Technology Co., Ltd.(江苏云河丰智能科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

↑