发表机构
Rutgers University(罗格斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过大规模审计发现ChatGPT对话中大量健康数据被隐式提取并合成永久性诊断特征,造成隐私风险,并提出设计指南以增强用户控制。
AI 中文摘要
随着对话式大语言模型深度融入日常生活,用户在常规交互中频繁披露敏感的个人健康信息。我们开展了一项大规模计算审计,分析了来自印度、尼日利亚、巴西和巴基斯坦的179,057段对话(N = 1,057),以评估ChatGPT中的个人健康披露和背景记忆合成情况。我们发现,21.31%的被审计对话包含个人健康数据,其中3.62%构成高至极端的隐私风险,涉及污名化疾病、直接标识符和精确位置。在评估ChatGPT的记忆条目时,我们揭示了企业宣传与系统行为之间的明显脱节:超过95%的档案条目是在没有明确用户提示或同意的情况下隐式提取的。此外,背景记忆合成选择性地将临时的、症状层面的披露压缩为永久性的诊断特征,剥离了情境完整性并放大了重新识别风险。最后,我们提出了社会技术设计指南,以在有状态AI系统中恢复用户自主权和基于同意的边界。
英文摘要
As conversational LLMs become deeply embedded in daily life, users frequently disclose sensitive personal health information during routine interactions. We present a large-scale computational audit analyzing 179,057 conversations across India, Nigeria, Brazil, and Pakistan (N = 1,057) to evaluate personal health disclosures and background memory synthesis in ChatGPT. We find that 21.31% of audited conversations contain personal health data, with 3.62% posing high-to-extreme privacy risks involving stigmatized conditions, direct identifiers, and precise locations. When evaluating the memory entries of ChatGPT, we uncover a stark disconnect between corporate framing and system behavior: over 95% of profile entries are implicitly extracted without explicit user prompts or consent. Furthermore, background memory synthesis selectively condenses temporary, symptom-level disclosures into permanent diagnostic traits, stripping contextual integrity and amplifying re-identification risks. We conclude with sociotechnical design guidelines to restore user agency and consent-driven boundaries in stateful AI systems.