arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

成员推断攻击对NLP文本分类器的实证评估:基于SST-2的基线研究

Empirical Evaluation of Membership Inference Attacks on NLP Text Classifiers: A Baseline Study on SST-2

William Novak, Muhammad Abusaqer

arXiv 2609.10935首次发表:更新:

发表机构

Minot State University(迈诺特州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文在SST-2上比较TF-IDF逻辑回归与DistilBERT的成员推断攻击脆弱性,发现两者均泄露成员信号,而轻量训练调整(如减少epoch)可在效用损失极小下改善隐私-效用权衡。

AI 中文摘要

成员推断攻击(MIAs)试图确定某个特定记录是否被用于训练模型,这是一种在自然语言处理(NLP)中至关重要的隐私风险,因为训练数据可能包含敏感的用户文本。本文针对GLUE SST-2情感数据集上的文本分类任务,提出了一个受控的成员推断脆弱性基准。在损失阈值MIA下,比较了TF-IDF + 逻辑回归流程和微调的DistilBERT分类器,效用通过开发集准确率和宏F1衡量。DistilBERT达到了0.9466的准确率和0.9460的宏F1,而逻辑回归分别为0.8756和0.8727,但两个模型都泄露了成员信号(攻击AUC分别为0.5615和0.5800)。测试了两种缓解措施。更强的正则化以明显的效用成本减少了逻辑回归的泄露,而将DistilBERT微调从3个epoch改为2个epoch,在准确率损失可忽略的情况下减少了泄露。轻量级的训练调整可以在不采用复杂防御的情况下改善隐私-效用权衡。

英文摘要

Membership inference attacks (MIAs) try to determine whether a specific record was used to train a model, a privacy risk that matters in natural language processing (NLP), where training data can contain sensitive user text. This paper presents a controlled benchmark of membership inference vulnerability for text classification on the GLUE SST-2 sentiment dataset. A TF-IDF + Logistic Regression pipeline and a fine-tuned DistilBERT classifier are compared under a loss-threshold MIA, with utility measured by development accuracy and macro F1. DistilBERT reached 0.9466 accuracy and 0.9460 macro F1 against 0.8756 and 0.8727 for Logistic Regression, yet both models leaked membership signal (Attack AUC 0.5615 and 0.5800, respectively). Two mitigations were tested. Stronger regularization reduced leakage for Logistic Regression at a visible utility cost, whereas fine-tuning DistilBERT for 2 epochs instead of 3 reduced leakage with negligible accuracy loss. Lightweight training adjustments can improve the privacy-utility trade-off without complex defenses.

CommentsPresented at the 58th Midwest Instruction and Computing Symposium (MICS 2026), Eau Claire, WI, March 27 to 28, 2026. 13 pages, 4 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑