arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

强多语言隐私标注:编码器速度下的实现

Strong Multilingual Privacy Tagging at Encoder Speed

Jonathan Graehl

arXiv 2609.38630首次发表:更新:

发表机构

RWS Language Weaver(RWS Language Weaver)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出一种多语言隐私命名实体标注器,通过微调编码器实现细粒度编辑,在七种语言上达到88.8的F1,显著优于现有方法,并支持低成本扩展和快速推理。

AI 中文摘要

隐私编辑必须在移除个人信息的同时保留文本中表达的关系。我们开发了一个多语言命名实体标注器,具有细粒度区分能力,支持多种编辑策略以及低成本学习额外区分的方法。我们使用仿射跨度标注头在35种语言的前沿模型标注上微调多语言编码器,通过覆盖感知掩码重放映射后的人工金标准,使未标注类型不被视为负例,并利用学习到的±1字符调整修复子词边界。在七种语言的1,283个人工金标准测试片段上,最佳实测编辑F1为88.8,而发布的GLiNER2为69.1(其任务排除了11种不可表示类型;若不排除则为68.8),适应新训练数据的GLiNER2为67.8,Microsoft Presidio为57.3,最佳发布的OpenAI隐私过滤器微调为35.8。增加约50,000条带标注的训练句子并增加人工金标准重放,将Ont3(我们包含31种类型的前沿标注NER评估,含1,201个开发片段)上的精确类型跨度F1从74.5提升至76.3。仅映射金标准重放即可将人工金标准F1提升10个百分点,且在前沿标注文本上无损失;边界调整在Ont3上增加了1.7个精确类型跨度F1点。在单个96GB GPU上运行的本地LLM作为提示标注器和冻结编码器表现不佳,编码速度比XLM-R推理慢30-95倍,提示标注在评估配置中约慢180-1,100倍。编码器架构的CPU吞吐量是GLiNER2的4.9倍。我们发布了代码、提示和训练配方,以及数据获取脚本和源代码链接。

英文摘要

Privacy redaction must remove personal information while preserving relationships expressed in text. We develop a multilingual named-entity tagger with fine-grained distinctions supporting varied redaction policies and methods for cheaply learning additional distinctions. We fine-tune a multilingual encoder with an affine span-tagging head on frontier-model annotations in 35 languages, replay mapped human gold with coverage-aware masking so unannotated types are not treated as negatives, and repair subword boundaries with a learned +/-1-character adjustment. On 1,283 human-gold test segments in seven languages, best measured redaction F1 is 88.8, against 69.1 for published GLiNER2 with 11 unrepresentable types excluded from its task (68.8 without that exemption), 67.8 for GLiNER2 adapted to the new training data, 57.3 for Microsoft Presidio and 35.8 for the best published OpenAI Privacy Filter fine-tune. Adding about 50,000 annotated training sentences and increasing human-gold replay improves exact typed-span F1 from 74.5 to 76.3 on Ont3, our 31-type frontier-annotated NER evaluation of 1,201 development segments. Mapped-gold replay alone raises human-gold F1 by ten points without loss on frontier-annotated text; boundary adjustment adds 1.7 exact typed-span F1 points on Ont3. Local LLMs fitting on a single 96-GB GPU underperformed as prompted annotators and frozen encoders, with encoding 30-95 times slower than XLM-R inference and prompted annotation roughly 180-1,100 times slower in the evaluated configurations. The encoder architecture delivers 4.9 times GLiNER2's CPU throughput. We release code, prompts and training recipes, with data-acquisition scripts and source links.

Comments46 pages, 23 figures. Includes supplementary appendices. Submitted to ACL Rolling Review, October 2026 cycle. v2: corrected citations and dataset licenses; human agreement reported over *all* multiply annotated TAB documents

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑