AI 中文总结
针对移动应用隐私政策与数据安全标签表述不一致问题,通过对6051款安卓应用大规模实证研究,利用基于大语言模型的框架和统一模式,测量一致性并引入风险评分,发现敏感类别受影响大,凸显披露机制差距,强调加强验证与提高透明度。
AI 中文摘要
随着移动应用的快速增长,用户数据隐私愈发受到关注。隐私政策描述应用如何收集和共享数据,而谷歌应用商店等平台提供数据安全标签来总结这些做法。但由于这些披露渠道是分开声明的,可能呈现出应用数据做法的不一致表述,给用户和监管机构带来不确定性。本研究对6051款安卓应用进行大规模实证研究,使用基于大语言模型的提取框架和统一模式,测量每个应用和每个类别的一致性,并引入强调高风险数据类型的敏感性加权风险评分。发现不一致对个人信息和设备标识符等敏感类别影响更大,共享披露的一致性低于收集披露。高隐私风险集中在与持续监控和通信相关的应用类别中。研究结果凸显了当前披露机制的结构性差距,强调在平台级隐私报告中需要更强的验证和更高的透明度。
英文摘要
With the rapid growth of mobile applications, user data privacy has become an increasing concern. While privacy policies describe how apps collect and share data, platforms such as Google Play provide Data Safety labels intended to summarize these practices. Because these disclosure channels are declared separately, they may present inconsistent representations of app data practices, creating uncertainty for users and regulators. In this work, we conducted a large-scale empirical study of disclosure consistency across 6,051 Android apps. Using an LLM-based extraction framework and a unified schema over 14 Google Play data categories and two operations (collection and sharing), we measure per-app and per-category consistency and introduce a sensitivity-weighted risk score that emphasizes high-risk data types. We find that misalignment disproportionately affects sensitive categories such as personal information and device identifiers, with sharing disclosures exhibiting lower consistency than collection disclosures. Elevated privac risk is concentrated in app categories associated with persistent monitoring and communication. Overall, our findings highlight structural gaps in current disclosure mechanisms and underscore the need for stronger verification and greater transparency in platform-level privacy reporting.
Journal refProceedings on Privacy Enhancing Technologies (PoPETs) 2026