arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PersianAnonymizer:评估基于LLM标注的训练,用于波斯语中基于NER的高效匿名化

PersianAnonymizer: Evaluating LLM-Labeled Training for Efficient NER-based Anonymization in Persian

Mohammad Hossein Shalchian, Mostafa Amiri, Amir Mahdi Sadeghzadeh

arXiv 2609.00958首次发表:更新:

发表机构

Sharif University of Technology; University of Tehran(谢里夫理工大学; 德黑兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对波斯语客户聊天匿名化,用三个指令调优LLM标注得到的语料库训练MatinaRoberta型NER模型,发现OSS_ZeroShot监督的NER在消费级GPU上可高效完成4万条消息的高质量匿名化。

AI 中文摘要

我们针对波斯语客户聊天的实用匿名化任务,通过基于LLM标注的监督训练紧凑型NER模型,并选择最佳标注器用于部署。我们对比三个指令调优的LLM:DeepSeek-V3-0324、GPT-OSS-120B和Qwen3-235B-A22B-Instruct-2507,使其在共享JSON协议下生成跨度标注,得到四个语料库(OSS_ZeroShot、Qwen_ZeroShot、Qwen_FewShot、DeepSeek_FewShot)。每个语料库训练一个基于MatinaRoberta的token分类器,并以token级别的精确率、召回率、F1值(整体及每类)进行评估。我们还报告标签覆盖召回率(LCR,即真实非O类token被预测为非O类的比例),并通过测试标注的token级韦恩图量化不同标注器间的行为差异。最后,我们对比LLM在H200节点上的测试集标注延迟,与训练后的NER在单块RTX 3090上的测试时标注延迟。结果显示,来自OSS_ZeroShot的监督能产生最强的宏F1值和LCR,而对应的NER模型在单块消费级GPU上约2分钟即可完成整个4万条消息测试集的标注,为波斯语工业数据的高质量、低成本匿名化确立了实用路径。

英文摘要

We target practical anonymization of Persian customer chats by training a compact NER model from LLM-labeled supervision and selecting the best labeler for deployment. We compare three instruction-tuned LLMs: DeepSeek-V3-0324, GPT-OSS-120B, and Qwen3-235B-A22B-Instruct-2507, to produce span annotations under a shared JSON protocol, yielding four corpora (OSS_ZeroShot, Qwen_ZeroShot, Qwen_FewShot, DeepSeek_FewShot). A MatinaRoberta-based token-classifier is trained per corpus and evaluated with token-level Precision/Recall/F1 (overall and per-class). We also report Label Coverage Recall (LCR), the proportion of gold non-O tokens predicted as non-O, and quantify cross-labeler behavior via a token-level Venn on test annotations. Finally, we contrast test-set annotation latency of the LLMs on H200 nodes with the trained NER's test-time labeling on a single RTX 3090. Results show that supervision from OSS_ZeroShot yields the strongest macro-F1 and LCR, while the resulting NER labels an entire 40K-message test set in approximately 2 minutes on one consumer GPU. This establishes a practical path to high-quality, low-cost anonymization for Persian industrial data.

Comments10 pages, 3 figures, 6 tables. Published at LREC 2026

Journal refProceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), pages 4497-4506, 11-16 May 2026. ELRA Language Resources Association (ELRA), 2026

DOI:10.63317/57u2ica9225o

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑