arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09164cs.AI

CIDER:用于隐私偏好对齐的上下文披露边界数据集

CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

Bingcan Guo, Eryue Xu, Jijie Zhou, Zhiping Zhang, Tianshi Li

首次发表
浏览论文内容

中文总结 AI 辅助

本文推出包含169名用户标注的CIDER数据集,用于评估LLM隐私偏好对齐,发现上下文个性化可提升预测准确率,GPT-5.4和Claude Sonnet 4.6表现更优。

中文摘要 AI 辅助

将大语言模型(LLM)与人类隐私偏好对齐,需要捕捉个体超出一般隐私规范的披露边界。然而,在现实场景中,获取这种细微偏好以评估对齐效果仍存在缺口。我们推出CIDER,这是一个包含169名用户的14850条人工标注的数据集,涵盖60个涉及违反隐私规范的信息共享的人际交流场景,形成1650个上下文披露边界集。每个边界代表真实用户在某一场景中,针对给定交流角色和AI介导条件,对9种共享变体做出的披露决策。我们设计了一项任务,要求模型根据历史边界预测用户的披露决策,任务中使用不同层级的上下文信息。在12种开源和专有模型中,仅使用6个历史示例,上下文个性化可将预测准确率提升最高达11.41个百分点。更大规模的模型如GPT-5.4(中等推理强度)和Claude Sonnet 4.6更擅长利用语义上下文理解用户特定、依赖场景的披露偏好,以实现更准确的预测,而较小规模的模型往往依赖基于披露粒度和可识别性的结构化启发式方法。个性化通常能提升预测准确率,但这种提升往往伴随模型的假阳性率和假阴性率出现不平衡变化,仅Claude Sonnet 4.6实现了两者的平衡提升。我们的研究结果揭示了推理时个性化在隐私偏好建模方面的潜力与局限,并将CIDER定位为推进个性化隐私对齐的资源。

英文摘要

Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general privacy norms. However, a gap remains in eliciting such nuanced preferences to evaluate alignment in realistic settings. We introduce CIDER, a dataset of 14,850 human annotations from 169 users, forming 1,650 contextual disclosure boundary sets across 60 interpersonal communication scenarios involving information sharing that violates privacy norms. Each boundary represents a real user's disclosure decisions over 9 sharing variants in a scenario, for a given communication role and AI-mediated condition. We formulate a task in which models predict a user's disclosure decision from historical boundaries, with varying levels of contextual information. Across 12 open and proprietary models, in-context personalization improves prediction accuracy by up to 11.41 percentage points using only 6 historical examples. Larger models such as GPT-5.4 (with medium reasoning effort) and Claude Sonnet 4.6 are better at leveraging semantic context to understand user-specific, context-dependent disclosure preferences for more accurate predictions, while smaller models tend to rely on structured heuristics based on disclosure granularity and identifiability. Personalization generally improves prediction accuracy, but the improvement is often accompanied by imbalanced shifts in false-positive and false-negative rates across models, with only Claude Sonnet 4.6 achieving balanced improvements in both. Our findings reveal both the promise and limitations of inference-time personalization for privacy preference modeling and position CIDER as a resource for advancing personalized privacy alignment.

发表机构

  • University of Washington(华盛顿大学)
  • UIUC(伊利诺伊大学厄巴纳-香槟分校)
  • Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑