发表机构
University of Edinburgh; Carnegie Mellon University(爱丁堡大学; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出PIA-Bench,首个评估大语言模型在真实隐私影响评估任务上的开放基准,基于73份结构化PIA(含451项风险与831项缓解措施)测试,发现现成LLMs能产生有意义评估,并呼吁改进领域工作流与基础设施。
AI 中文摘要
隐私影响评估(PIA)是机构在系统部署前主动识别隐私风险并制定缓解策略的关键工具。尽管在监管和机构环境中被强制要求,执行PIA需要广泛的隐私和技术专业知识,这对缺乏此类资源的团队构成了特别挑战。先前的研究显示了利用大语言模型(LLMs)协助从业者隐私决策的潜力,但关于LLMs能多准确、多可靠地自动化PIA,我们知之甚少。为此,我们开发了PIA-Bench,这是首个用于在真实世界PIA上评估LLMs的开放基准。我们首先审计了美国联邦机构发布的499份专家撰写的PIA,并整理了73份结构化的PIA,共包含451项隐私风险和831项缓解措施,以评估LLMs评估复杂系统的隐私风险并提出缓解措施的能力。我们的结果表明,现成的LLMs能产生有意义的评估,并指出了未来改进的方向。最后,我们呼吁改进面向LLM智能体的领域特定工作流程,开发负责任的LLM基础设施,并为PIA设计新的质量标准。
英文摘要
Privacy impact assessment (PIA) is a critical instrument for institutions to proactively identify privacy risks and develop mitigation strategies before system deployment. While mandated across regulatory and institutional contexts, executing PIA requires extensive privacy and technical expertise, posing a particular challenge for teams without access to such resources. Prior work shows the potential of leveraging large language models (LLMs) to assist practitioners' privacy decisions, but little is known about how accurately and reliably LLMs can automate PIA. To this end, we develop PIA-Bench, the first open benchmark for evaluating LLMs on real-world PIAs. We first audited 499 expert-authored PIAs published by US federal agencies and curated 73 structured PIAs, comprising a total of 451 privacy risk and 831 mitigation items, to evaluate LLMs' ability to assess privacy risks and propose mitigations of complex systems. Our results show that off-the-shelf LLMs produce meaningful assessments and identify avenues for future improvement. Finally, we call for improving domain-specific workflows for LLM agents, developing accountable LLM infrastructure, and designing new quality standards for PIAs.