发表机构
Varitas(瓦里塔斯公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究定义“隐私清洗”,通过四阶段流程结合三模型LLM小组投票检测隐私政策内部矛盾,发现11年语料库中第三方共享矛盾占多数,还分析了流行率及配置对结果的影响。
AI 中文摘要
隐私政策可能包含内部矛盾,即政策中某处记录的实践会破坏另一处的承诺。我们将这一现象定义为“隐私清洗”,并通过四阶段流程对其进行研究:语句提取、兼容性过滤与自然语言推理筛选、多模型判断验证,以及主题分析,矛盾由一个包含三个大型语言模型(LLM)的小组通过多数投票确认。该流程应用于两个网站隐私政策语料库:2026年收集的123份(OPPT)和2015年收集的115份(OPP-115),结果发现这两个间隔11年的语料库中存在相同的类别模式,第三方共享矛盾是每次主要运行中已确认案例的多数,这与政策构成的结构因素一致,而非必然是故意欺骗。在OPPT公司中,至少存在一个经小组确认的矛盾的比例为12.2%(15/123;若排除遗留对则为9.8%);在OPP-115公司中,这一比例为36.5%(42/115)。七个月后进行的稳定性重运行采用了完全分离的配置(新的提取模型、来自两个语料库之外的三家中国供应商的判断者、匹配的过滤器、无判断者-提交物相似度阈值),在原始协议下重现了OPPT的流行率(13.0% vs. 12.2%),发现亚阈值对的确认率与上述比例处于同一数量级(使流行率分别升至20.3%和40.9%),且显示第三方多数对小组敏感,而相同类别对的重现则不敏感。所有数据存在两个注意事项:小组的裁决未通过人类专家判断验证,因此精确率未知,流行率数字为下界;两次主要运行使用了不同的过滤器配置,因此它们的流行率差异不能解释为语料库或时代效应(匹配重运行将差距缩小至约两倍,但未消除)。
英文摘要
Privacy policies may contain internal contradictions in which commitments are undermined by practices documented elsewhere in the same policy. We operationalize this phenomenon, privacy washing, through a four-stage pipeline: statement extraction, compatibility filtering and natural language inference screening, multi-model judge verification, and thematic analysis, with contradictions confirmed by majority vote of a three-model LLM panel. Applied to two corpora of website privacy policies, 123 collected in 2026 (OPPT) and 115 collected in 2015 (OPP-115), the pipeline finds the same category patterns recurring across the 11-year gap, with third-party sharing contradictions the majority of confirmed cases in each primary run, consistent with structural factors in policy composition rather than necessarily intentional deception. At least one panel-confirmed contradiction appears in 12.2% of OPPT companies (15/123; 9.8% excluding legacy pairs) and 36.5% of OPP-115 companies (42/115). A stability re-run seven months later, with a fully separated configuration (new extraction models, judges from three Chinese providers absent from both corpora, matched filters, no judge-submission similarity threshold), reproduces the OPPT prevalence under the original protocol (13.0% vs. 12.2%), finds sub-threshold pairs confirm at rates of the same order as those above (raising prevalence to 20.3% and 40.9%), and shows the third-party majority is panel-sensitive while the recurrence of the same category pairs is not. Two caveats govern all figures: panel verdicts are not validated against human expert judgment, so precision is unknown and prevalence figures are lower bounds; and the two primary runs used different filter configurations, so their prevalence difference is not interpretable as a corpus or era effect (the matched re-run reduces the gap to roughly twofold but does not eliminate it).