Redakto——面向大语言模型的隐身标签
Redakto - The Incognito Tab for LLMs
查看机构详情
- Berlin University of Applied Sciences(柏林应用科学大学)
- Einstein Center Digital Future(爱因斯坦数字未来中心)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
Redakto是一款开源匿名化工具,可通过网页应用、REST API等供用户使用,能移除文本中PII,其匿名化后的文本效用与原文相当,可助力LLM使用时的隐私保护。
中文摘要 AI 辅助
大语言模型(LLM)正越来越多地应用于日常场景。对于LLM或通用人工智能(AI)而言,一个核心挑战是确保使用时的隐私性,即从输入LLM的所有文本中移除个人可识别信息(PII)。随着欧盟新法规的出台,这些挑战变得愈发紧迫。欧盟各国在LLM使用方面的隐私不确定性,可能成为阻碍创新速度及研究向应用转化的重要因素。本文提出Redakto,一款可在将文本输入LLM或其他下游文本处理任务前用于匿名化的工具。它提供了处于先进水平的功能,既支持PII的编辑,也可用于假名化。这些功能可通过Redakto网页应用供终端用户轻松使用,也可通过REST API和模型上下文协议(MCP)钩子供开发者和研究者使用。该实现完全开源,仅需适度计算资源,可轻松部署在本地硬件上。与先前工作不同,为更好评估匿名化文本的质量,我们针对法律和医疗领域的文本,从隐私性和编辑后文本的效用两方面开展了广泛的实证评估。实证结果表明,采用不同编辑策略匿名化后的文本,其效用得分与原始文本相当,这表明使用Redakto进行匿名化,可在所探索的LLM任务中使用,且不会对任务产生实质性负面影响。
英文摘要
Large Language Models (LLMs) are being increasingly used in everyday applications. A major challenge in the context of LLMs or Artificial Intelligence (AI) in general is to ensure privacy when using them, meaning that personally identifiable information (PII) is removed from any text that enters an LLM. These challenges have become more urgent with novel EU legislation. Uncertainty around LLM usage with respect to privacy concerns in EU countries can be a major blocker for the speed of innovation and transfer from research to applications. Here we present \textbf{Redakto}, a tool that can be used for anonymizing text prior to feeding it to an LLM or other downstream text processing. We provide state-of-the-art functionalities for both redaction of PII but also when used for pseudonymization. These functionalities are exposed such that they can easily be used by end-users, through the Redakto web application, and by developers and researchers, via REST APIs and model context protocol (MCP) hooks. The implementation is fully open source, requires modest compute resources, and can be readily deployed on local hardware. In contrast to prior work and in order to better assess the quality of the anonymized texts, we conduct extensive empirical evaluations on textual data from legal and medical domain with respect to both privacy and utility of the redacted texts. Our empirical results demonstrate that the texts anonymized with different redaction strategies achieve utility scores on par with the original texts, suggesting that anonymization with Redakto can be used for LLM tasks without substantial negative impact for the tasks we explored.