arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

阅读并非泄露:对公共足迹推理暴露的本地、可审计测量与削减

Reading Is Not Leaking: Local, Auditable Measurement and Reduction of Inference Exposure from Public Footprints

Mahmudul Faisal Al Ameen

arXiv 2609.32565首次发表:更新:

发表机构

Independent researcher France; Independent researcher(; )

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出本地可审计框架,分离阅读与泄露以测量推理暴露,用106M编码器提供可验证证据,并以低成本改写防御削减泄露。

AI 中文摘要

任何拥有公共足迹的人都会泄露从未被陈述过的事实,而语言模型使得这种推理变得廉价。我们提出一个框架,用于测量和减少这种推理暴露,该框架在所有者自己的CPU上运行,分析时无需语言模型,并已在组织和个体上实例化。它始于一个测量结果:将推理系统针对目标私有真相进行评分,混淆了系统阅读记录的能力与记录泄露的程度。在针对十六家合成企业的128题工具上,几乎一半的问题从未被六位读者中的任何一位正确回答,其中四位是语言模型,而多数类猜测占据了每位读者得分的大部分。因此,我们将阅读准确性与泄露率分离,并引入一种注入协议,创建具有已知支持的单元。我们的分析器结合规则、统计求解器和一个从头训练的106M参数编码器,该编码器标记逐字证据且从不生成文本;每个答案都带有分级证书,其记录的证明可重放。其认证答案在93%的已解决案例中正确,而语言模型基于引用的答案正确率为49-73%,这些引用与答案同时生成而非推导出答案;当记录中包含平实散文文章时,其70%的带证据答案依赖于能确立它们的证据,而模型的这一比例为18-56%。一种约束性防御将每个事实的载体改写为真实但更粗略的陈述,以比删除低40%的编辑成本,将所有单载体事实隐藏于四位语言模型对手。在十六个合成个体上,猜测项更大,且一个无语言模型的诱饵规划器将其目标估计器的正确答案减半,而不会转移至第二个估计器。

英文摘要

Anyone with a public footprint leaks facts that were never stated, and language models make the inference cheap. We present a framework for measuring and reducing this inference exposure that runs on the owner's own CPU with no language model at analysis time, instantiated on organisations and on individuals. It starts from a measurement result: scoring an inference system against the target's private truth conflates how well the system reads the record with how much the record leaks. On a 128-question instrument over sixteen synthetic firms, almost half of the questions are never answered correctly by any of six readers, four of them language models, and a majority-class guess accounts for most of every reader's score. We therefore separate reading accuracy from leakage rate and introduce an injection protocol that creates cells with known support. Our analyser combines rules, statistical solvers and a 106M-parameter encoder trained from scratch that marks verbatim evidence and never generates text; every answer carries a graded certificate whose recorded proof replays. Its certified answers are correct in 93% of resolved cases, against 49-73% for the language models' quote-backed answers, whose citations are produced alongside the answer rather than deriving it; with plain-prose articles in the record, 70% of its evidence-bearing answers rest on evidence that establishes them, against 18-56% for the models. A constrained defence that rewrites each fact's carrier as a true but coarser statement hides every single-carrier fact from four language-model adversaries at 40% lower edit cost than deletion. On sixteen synthetic people the guessing term is larger still, and a decoy planner with no language model halves the correct answers of the estimator it targets without transferring to a second.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑