arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

隐私友好的群体确定:密封的、独立于CSP的浏览器内机器学习推断专业群体,用于无身份广告

Privacy-Friendly Cohort Determination: Sealed, CSP-Independent In-Browser ML Inference of Professional Segments for Identity-Less Advertising

Om Shankar Tiwari, Navnit Shukla, Guanyu Wang, Akshay Jain

arXiv 2609.36153首次发表:更新:

发表机构

Google; Snowflake Inc.; TikTok; LinkedIn(谷歌; 雪花公司; 抖音国际版; 领英)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对B2B广告中的隐私问题,提出SIF框架,在设备端通过WebAssembly推断职业群体,采用本地差分隐私和k-匿名规则,仅发送分类编码标签,实现无身份广告投放。

AI 中文摘要

B2B广告针对观看者的职业属性(雇主规模和行业、职能、资历),并通过跨站点匹配身份来获取这些属性。Safari和Firefox阻止第三方Cookie,Google在2025年退役了隐私沙盒群体API,而基于反向IP的企业画像在远程工作环境下逐渐失效。我们提出SIF(密封推断框架),它在设备上推断粗略的职业群体,并仅将本地差分隐私的、基于分类法编码的标签发送到OpenRTB竞价流中,不涉及跨站点标识符。它依赖于我们精确阐述的Web平台的一个属性:导航的跨源iframe是第三方代码获得其控制的策略的唯一方式,因此即使发布者的CSP禁止,推断也在WebAssembly中运行,并且一个以default-src 'none'提供的嵌套worker使模型无法访问网络。即使模型是恶意的,每个站点每周最多泄漏约5比特信息。标签通过一个以发布者第一方标识符为键的记忆化k元随机响应传递,这提供了ε-本地差分隐私,抵御平均化攻击,并且链接请求不比已发送的标识符更好。一个组织条件的k-匿名规则抑制单元格,在企业网络上比在家庭中更严格。群体通过OpenRTB this http URL在LinkedIn对齐的分类法中传输,归因使用LinkedIn的点击范围的li_fat_id,而不桥接身份。我们报告了对7,969个顶级站点和431个B2B发布者的CSP部署的爬取、Heavy-Ad预算、闭式隐私-效用权衡、一个重识别模拟,以及评估哪些属性是可预测的:公司类型和规模是可预测的,资历在很大程度上不可预测。设备端是一个设计属性,而不是同意豁免。

英文摘要

B2B advertising targets a viewer's professional attributes (employer size and industry, function, seniority) and has obtained them by matching identities across sites. Safari and Firefox block third-party cookies, Google retired the Privacy Sandbox cohort APIs in 2025, and reverse-IP firmographics decay under remote work. We present SIF (Sealed Inference Frame), which infers coarse professional cohorts on the device and emits only a locally differentially private, taxonomy-coded label into the OpenRTB bid stream, with no cross-site identifier. It rests on a property of the web platform we make precise: a navigated cross-origin iframe is the only way third-party code obtains a policy it controls, so inference runs in WebAssembly even where the publisher's CSP forbids it, and a nested worker served with default-src 'none' gives the model no network. Even a malicious model leaks at most about 5 bits per site per week. Labels pass through a memoised k-ary randomised response keyed to the publisher's first-party identifier, which gives $\varepsilon$-local differential privacy, defeats averaging, and links requests no better than the identifier already sent. An org-conditional k-anonymity rule suppresses cells, more strictly on corporate networks than at home. Cohorts ride OpenRTB user.data in a LinkedIn-aligned taxonomy, and attribution uses LinkedIn's click-scoped li_fat_id without bridging identities. We report a crawl of CSP deployment on 7,969 top sites and 431 B2B publishers, Heavy-Ad budgets, closed-form privacy-utility trade-offs, a re-identification simulation, and an assessment of which attributes are predictable at all: company type and size are, seniority largely is not. On-device is a design property, not a consent exemption.

Comments11 pages, 5 figures, 2 tables. Ancillary code, crawl data and full proofs are included

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑