发表机构
Petscraft, Inria; Université Paris-Saclay; INSA CVL; Université d’Orléans; LIFO; ÉTS Montréal(Petscraft,法国国家信息与自动化研究所; 巴黎萨克雷大学; 中央-卢瓦尔河谷国立应用科学学院; 奥尔良大学; LIFO实验室; 蒙特利尔高等工程技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出PrivacySkills框架,通过55个合成任务评估隐私指导(系统指令、技能标签或两者结合)对LLM智能体信息源选择的影响,发现结合两种指导可将机密访问率减半,建议将隐私注释纳入技能规范。
AI 中文摘要
尽管先前的研究已记录了LLM智能体中的隐私失败,但隐私指导的呈现方式如何影响其信息源选择仍不清楚。我们引入了PrivacySkills,一个受控框架,用于评估智能体如何在提供相同任务相关价值的获取路径中进行选择:查阅公开可用的个人信息、访问机密来源或与用户交互。该评估框架包含55个合成任务,涵盖11类个人信息,并附带169个描述可用获取路径的相关技能。我们通过系统级指令、技能级元数据标签或两者结合来考虑隐私指导。此外,我们分别变化用户可用性和紧迫性框架。在用户可用且无隐私指导的情况下,五个开放权重模型在有效运行中平均有30%的几率访问机密来源,尽管存在足够的替代方案。当用户不可用时,这一比例上升至45%,而紧迫性框架没有可检测的影响。仅系统级隐私指令对机密访问的影响有限,而技能级侵入性标签产生适度降低(平均24%),但两者结合可将机密访问大致减半。我们的发现激励将隐私注释纳入技能规范,并评估其与系统级指令相结合的有效性。
英文摘要
While prior work has documented privacy failures in LLM agents, it remains unclear how the presentation of privacy guidance influences their choice of information sources. We introduce PrivacySkills, a controlled framework for evaluating how agents choose among acquisition pathways that provide the same task-relevant value: consulting publicly available personal information, accessing confidential sources, or interacting with the user. The evaluation framework comprises 55 synthetic tasks spanning 11 categories of personal information, with 169 associated skills that describe the available acquisition pathways. We consider privacy guidance through system-level instructions, skill-level metadata labels, or both. Separately, we vary user availability and urgency framing. With users available and no privacy guidance, agents access confidential sources in 30% of valid runs on average across five open-weight models, despite sufficient alternatives. This rate increases to 45% when users are unavailable, whereas urgency framing has no detectable effect. System-level privacy instructions alone have limited effects on confidential access, while skill-level intrusiveness labels produce a modest reduction (24% on average), but combining the two roughly halves confidential access. Our findings motivate incorporating privacy annotations into skill specifications and evaluating their effectiveness alongside system-level instructions.