arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过AI介导的智能手机数据分析导致的隐私泄露

Privacy Leakage Through AI-mediated Analysis of Smartphone Data

Sarah Radway, Zoe Robert, Matthew Soto, Julianna Cimillo, Sebastian Diaz, Meg Marco, James Mickens

arXiv 2609.28537首次发表:更新:

发表机构

Harvard University; Carnegie Mellon University(哈佛大学; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能手机应用数据收集引发的隐私风险,构建基于LLM的Priva-See推理系统,经465人研究证明其能仅凭部分数据做出隐私侵犯推断,并据此提出改进系统同意机制的方案。

AI 中文摘要

在过去三十年中,在线广告行业构建了一个大规模的数据收集生态系统,其目标是跟踪用户的在线活动,以推断其人口统计信息和兴趣。传统上,该生态系统依赖于对高度结构化的文本数据(如用户IP地址、GPS坐标、电子商务购买历史和访问的URL)的整理和分析。然而,最近的机器学习模型不仅能够解析结构化文本,还能解析多媒体文件和非结构化文本输入——这意味着用户的照片、视频、收件箱和日历现在都适合进行自动化分析。在智能手机应用的背景下,隐私风险尤为严重。用户的手机已经自然地成为敏感用户信息的汇集点,但用户可能并不理解,允许某个应用(例如)访问用户的照片,不仅仅赋予该应用访问照片字节的权限:该应用还获得了通过照片对用户进行推断的能力。为了探索这些隐私风险,我们构建了Priva-See,一个基于LLM的推理系统,用于分析应用收集的用户数据;Priva-See反映了我们对现实广告技术公司如何利用机器学习构建用户画像的最佳理解。通过一项经IRB批准的用户研究,465名参与者在他们的手机上部署了Priva-See;尽管Priva-See只能访问用户数据的一个子集,但它仍做出了侵犯隐私的推断。我们发现,这一经历显著影响了参与者未来共享权限数据的意愿。基于观察到的隐私侵犯行为,我们建议对智能手机操作系统如何收集用户数据访问同意进行更改,以更好地告知用户下游数据使用能力。

英文摘要

Over the past thirty years, the online advertising industry built a large-scale data collection ecosystem, with the goal of tracking a user's online activity to infer their demographics and interests. Traditionally, the ecosystem relied upon the collation and analysis of highly-structured text data like user IP addresses, GPS coordinates, e-commerce purchase histories, and visited URLs. However, recent ML models can parse not only structured text, but also multimedia files and unstructured text inputs---meaning a user's photos, videos, inboxes, and calendars are now ripe for automated analysis. The privacy risks are particularly acute in the context of smartphone apps. A user's phone already acts as a natural collation point for sensitive user information, but users may not understand that permitting an app to, for example, access a user's photo does not just give the app access to the bytes in the photo: the app also receives access to inferences about the user that are enabled by the photo. To explore these privacy risks, we built Priva-See, an LLM-based inference system for app-collected user data; Priva-See reflects our best understanding of how real-life adtech companies would leverage machine learning to build user profiles. Through an IRB-approved user study, 465 participants deployed Priva-See on their phones; Priva-See made privacy-invasive inferences despite having access to only a subset of a user's data. We see the experience significantly impacted participant willingness to share permissions data moving forward. Based on the observed privacy violations, we suggest changes to how smartphone OSes should gather user consent for data access, to better inform users about downstream data usage capability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑