发表机构
Technical University of Munich; Munich Center for Machine Learning; Munich Data Science Institute(慕尼黑工业大学; 慕尼黑机器学习中心; 慕尼黑数据科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出整体模块化基准测试平台PrivBench,用于统一评估文本到文本私有化,具备可扩展性、用户友好性,通过实时排行榜推动竞争,免费公开可用。
AI 中文摘要
自然语言处理方法已为隐私领域带来新的解决方案和进展,尤其是在文本到文本私有化这一子领域,其目标是将敏感输入文本转换为私有化输出,理想情况下需掩盖直接或间接可识别的信息或其他私人信息。然而,文本到文本私有化的评估并非易事,现有文献已采用大量技术和指标来量化私有化方法的隐私保护能力。为统一文本到文本私有化的评估,我们推出PrivBench——一个面向从事文本私有化研究的研究者和从业者的整体模块化基准测试平台。PrivBench的整体性体现在它会根据一系列既定的理想特性(desiderata)对私有化进行评估,这些特性被组织成多个模块。PrivBench不仅是模块化的,还具有可扩展性,支持未来的更新和基准版本迭代。该平台以用户为中心,通过实时评估和公开实时排行榜推动竞争,可免费使用,公开访问地址为this https URL。
英文摘要
Natural Language Processing methods have enabled novel solutions and advances in the field of privacy, particularly in the sub-domain of text-to-text privatization, where the goal is to transform a sensitive input text into a privatized output by ideally masking (in)directly identifiable or otherwise private information. The evaluation of text-to-text privatization, however, is not straightforward, and the extant literature has utilized a myriad of techniques and metrics to quantify the privacy-preserving capabilities of privatization methods. Seeking to unify the evaluation of text-to-text privatization, we introduce PrivBench, a holistic and modular benchmarking platform for researchers and practitioners working on text privatization. PrivBench is holistic in that it evaluates privatization on a series of defined desiderata, which are structured into modules. PrivBench is not only modular but also extensible, allowing for future updates and benchmark versions. PrivBench is user-centered and promotes competition via real-time evaluation and a live public leaderboard. The platform is free to use and openly accessible at https://privbench.com/.
Comments23 pages, 5 figures, 3 tables, accepted to EMNLP 2026 System Demonstrations