GRASP:用组相对策略优化增强语言模型匿名化器
GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy Optimization
查看机构详情
- UC Santa Barbara(加州大学圣巴巴拉分校)
- UC Los Angeles(加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究提出GRASP方法,用组相对策略优化增强设备端小型语言模型匿名化器,在隐私-效用权衡、对抗匿名化防御及运行成本上均优于DPO蒸馏基线。
中文摘要 AI 辅助
大型语言模型可从普通文本中推断年龄、位置、职业等敏感个人属性,将日常写作转化为隐私风险。对抗匿名化通过用强大语言模型重写文本(该模型同时充当攻击者)来防御,但它推理时需要强大模型,因此会将私人文本发送给第三方,而这正是匿名化本应防止的暴露。近期工作通过监督微调与直接偏好优化(DPO)将此行为蒸馏为小型设备端模型,但DPO仅模仿教师的离线选择,从未直接优化我们关心的隐私-效用目标。我们提出GRASP(通过自求精策略优化实现的组相对匿名化,Group-Relative Anonymization via Self-refinement Policy-optimization),其用组相对策略优化在线增强本地匿名化器。单个小型模型同时充当匿名化器、攻击者与效用评判器,针对自生成的奖励进行训练,该奖励在隐藏属性的同时保留语义,设计上可防范奖励黑客攻击。在Llama-3.1-8B上训练后,\textbf{GRASP}在三个独立LLM评判器上均优于DPO蒸馏基线的隐私-效用权衡;针对Gemini 2.5 Flash、Claude等前沿模型驱动的对抗匿名化,它实现了相当或更优的整体权衡,同时移除了多得多的私人信息,且完全在设备端运行,成本约为GPT-4o教师模型的1%。
英文摘要
Large language models can infer sensitive personal attributes, such as age, location, and occupation, from ordinary text, turning everyday writing into a privacy risk. Adversarial anonymization defends against this by rewriting a text with a capable language model that also plays the attacker, but it needs a powerful model at inference time and thus sends private text to a third party, the very exposure anonymization should prevent. Recent work distills this behavior into a small on-device model using supervised fine-tuning and direct preference optimization (DPO), but DPO only imitates the teacher's offline choices and never directly optimizes the privacy--utility objective we care about. We introduce \textbf{GRASP} (\textbf{G}roup-\textbf{R}elative \textbf{A}nonymization via \textbf{S}elf-refinement \textbf{P}olicy-optimization), which reinforces the local anonymizer online with Group Relative Policy Optimization. A single small model acts as anonymizer, adversary, and utility judge, trained against a self-generated reward that hides attributes while preserving meaning, with a design that guards against reward hacking. Trained on Llama-3.1-8B, \ours{} improves the privacy--utility trade-off over the DPO-distilled baseline, consistently across three independent LLM judges. Against adversarial anonymization driven by frontier models such as Gemini~2.5~Flash and Claude, it achieves a comparable or better overall trade-off while removing substantially more private information, and it runs entirely on-device at roughly $1\%$ of the GPT-4o teacher's cost.