发表机构
University of South Florida; IBM T.J. Watson Research Center(佛罗里达州立大学; IBM 汤普森研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文系统评估现成大语言模型能否取代专业隐私分析工具,研究六个具三种主要功能及三个中间任务的工具,对比两个先进大语言模型与工具性能,发现大语言模型在隐私政策和法规分析功能上能匹配或超越现有工具。
AI 中文摘要
大语言模型(LLMs)的出现显著改变了隐私政策和数据合规性分析研究,使以前需要特定领域工具的任务得以实现。然而,LLMs在多大程度上能真正复制先前工作提供的多样功能、方法和分析尚不清楚。本文首次系统评估现成的LLMs能否取代专业隐私分析工具。研究了涵盖三种主要功能(矛盾检测、法规合规分析、隐私政策总结与聚合)及三个中间任务(使用元组进行结构化数据提取、语义角色标注(SRL)和手动隐私政策标注)的六个代表性工具。通过直接促使模型在10个隐私政策的自定义数据集上执行相应功能和任务,比较了两个最先进的LLMs(不同配置下的GPT-5.2和Gemini-2.5)与工具的性能,以评估现成模型能否在无需进一步工程或特定领域训练的情况下产生工具特定功能。结果表明,LLMs在各项功能上始终能匹配或超越现有工具。在第一方收集实体的手动标注中,LLMs平均精确率达81.8%,召回率达70.9%;在第三方共享实体标注中,与OPP-115数据集相比,平均精确率为91.4%,召回率为70.8%。总体而言,研究结果表明LLMs能有效执行隐私政策和法规分析中以前需要专业工具的广泛功能和任务。
英文摘要
The advent of LLMs has significantly changed the research on privacy policy and data compliance analysis by enabling tasks that previously required specialized, domain-specific tools. However, it remains unclear to what extent LLMs can truly replicate the diverse functionalities, and the wide range of methodologies and analysis offered by prior work. In this paper, we conduct the first systematic evaluation of whether off-the-shelf LLMs can replace specialized privacy analysis tools. We study six representative tools spanning three major functionalities: contradiction detection, regulatory compliance analysis, and privacy policy summarization and aggregation, and across three intermediate tasks: structured data extraction using tuples, Semantic Role Labeling (SRL) and manual privacy policy labeling. We compare the performance of two state-of-the-art LLMs (GPT-5.2 and Gemini-2.5 in various configurations) against the tools by directly prompting the models to perform corresponding functionalities and tasks on a custom dataset of 10 privacy policies, allowing us to assess whether off-the-shelf models can produce tool-specific functionalities without further engineering or domain-specific training, major limitations in prior work. Our results show that LLMs consistently match or exceed the capabilities of existing tools across the functionalities. In manual labeling of first-party collection entities, LLMs achieved an average precision of 81.8% and recall of 70.9%, while for labeling of third-party sharing entities, they achieved an average precision of 91.4% and recall of 70.8% compared to the OPP-115 dataset. Overall, our findings indicate that LLMs can effectively perform a broad range of functionalities and tasks in privacy policy and regulation analysis that previously required specialized tools.