arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TextCloak:基于强化学习的不可学习文本防御未经授权的大语言模型利用

TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text

Chengshuai Zhao, Pingchuan Ma, Dawei Li, Bohan Jiang, Zhiyuan Yu, Zhen Tan, Huan Liu

arXiv 2607.28862首次发表:更新:

AI 中文总结

本研究提出TextCloak框架,采用强化学习与GRPO-UE技术生成不可学习文本,可削弱大语言模型的未经授权微调,兼具实用性与鲁棒性。

AI 中文摘要

大语言模型(LLMs)的快速发展推动了各类语言任务的显著进步,同时也引发了对未经授权的数据利用与隐私泄露的日益担忧。不可学习示例(UEs)通过向数据中引入精心设计的扰动,使在其上训练的模型效用下降,提供了一种颇具前景的防御思路。然而,现有的文本保护方法主要针对判别式语言模型的分类任务(如情感分析)设计,且通常依赖注入类别特定的语言线索,这限制了它们在LLMs开放生成场景中的有效性。本研究提出TextCloak,一种用于保护文本数据免受未经授权LLM利用的强化学习驱动框架。TextCloak采用生成策略,将批量干净文本转换为不可学习示例,同时保留语义保真度与语言自然度。为优化该策略,我们引入GRPO-UE,其基于生成的不可学习文本对微调后的代理LLMs造成的下游性能下降来给予奖励,并通过组相对策略优化更新生成器参数。这种双层优化使生成器能够发现超越类别特定线索的可泛化保护模式。在六个公开数据集和九个最先进LLMs上开展的综合实验表明,TextCloak在持续削弱未经授权微调的同时,保持了文本在合法使用中的效用。进一步分析证实其在模型架构、训练配置及自适应攻击间具有可迁移性与鲁棒性,凸显了其作为防御未经授权LLM利用的实用防御手段的广泛适用性。

英文摘要

The rapid development of Large Language Models (LLMs) has led to significant advances across a wide range of language tasks, while simultaneously raising growing concerns about unauthorized data exploitation and privacy leakage. Unlearnable examples (UEs) offer a promising defense by introducing carefully designed perturbations into data such that models trained on them exhibit degraded utility. However, existing methods for text protection are primarily designed for classification tasks (e.g., sentiment analysis) in discriminative language models and often rely on injecting class-specific linguistic cues, which limits their effectiveness in the open-ended generation settings of LLMs. In this work, we propose TextCloak, an RL-driven framework for protecting textual data against unauthorized LLM exploitation. TextCloak employs a generative policy that transforms batches of clean text into unlearnable examples while preserving semantic fidelity and linguistic naturalness. To optimize the policy, we introduce GRPO-UE, which rewards generated unlearnable text based on the downstream degradation they induce in fine-tuned surrogate LLMs and updates the generator parameters via group-relative policy optimization. This bi-level optimization enables the generator to discover generalizable protective patterns beyond class-specific cues. Comprehensive experiments on six publicly available datasets and nine state-of-the-art LLMs demonstrate that TextCloak consistently impairs unauthorized fine-tuning while maintaining text utility for legitimate use. Further analyses establish its transferability and robustness across model architectures, training configurations, and adaptive attacks, highlighting its broad applicability as a practical defense against unauthorized LLM exploitation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑