LLM 去品牌化:在保留通用功能的同时消除商业标识
LLM unbranding: Erasing Commercial Identity while Preserving Generic Utility
- Jagiellonian University(雅盖隆大学)
- IDEAS Research Institute(IDEAS 研究所)
- Heinrich Heine Universität Düsseldorf(杜塞尔多夫海因里希·海涅大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对 LLM 在文本中生成品牌描述带来的风险,本文提出 LLM 去品牌化任务,并推出 MUTE 推理时方法,通过迭代优化系统指令消除品牌泄漏,同时保持模型通用能力。
AI中文摘要:
在图像生成领域,建立去品牌化作为一种关键实践以防止视觉标志获得负面含义已成为标准做法。大型语言模型(LLM)如今面临一个平行且新兴的挑战。这些模型经常在各种上下文中生成品牌描述。这种频繁性带来了显著风险,如商标淡化、错误归因和品牌诽谤。为此,我们正式定义了 LLM 去品牌化这一新任务。我们特别处理在文本输出中管理商业外观这一复杂挑战。这涉及中和定义品牌身份的特征性语言、口号和风格标记。关键的是,这些元素不如显式视觉标志那么明显。为了对这项任务进行基准测试,我们引入了一个全面的评估数据集,涵盖了多个商业领域的知名品牌。我们使用该基准严格评估了现有的最先进机器遗忘模型。该评估揭示了它们在选择性文本去品牌化方面的局限性。最后,我们提出了 MUTE,一种新颖的推理时方法,能够有效中和文本商业外观,同时保留 LLM 的通用能力和实用性。通过利用迭代细化循环,MUTE 系统地优化系统指令,以安全地消除品牌泄漏,而无需脆弱的参数更新。代码和数据集:LLM 去品牌化的评估数据集和代码可在该 https URL 获取。MUTE 的实现可在该 https URL 获取。
英文摘要:
Establishing unbranding as a critical practice to prevent visual logos from acquiring negative connotations is standard in image generation. Large Language Models (LLMs) now face a parallel and emerging challenge. These models frequently generate brand descriptions within diverse contexts. This frequency introduces significant risks, such as trademark dilution, false attribution, and brand defamation. In response, we formally define the novel task of LLM Unbranding. We specifically address the complex challenge of managing trade dress within textual outputs. This involves neutralizing characteristic language, slogans, and stylistic markers that define brand identity. Crucially, these elements are less evident than explicit visual logos. To benchmark this task, we introduce a comprehensive evaluation dataset incorporating prominent brands from multiple commercial domains. We rigorously evaluate existing state-of-the-art machine unlearning models using this benchmark. This evaluation identifies their limitations in selective textual unbranding. Finally, we propose MUTE, a novel inference-time method that effectively neutralizes textual trade dress while preserving the LLM's general capabilities and utility. By leveraging an iterative refinement loop, MUTE systematically optimizes system instructions to safely eliminate brand leakage without requiring fragile parameter updates. Code and dataset: The evaluation dataset and code for LLM Unbranding are available at https://github.com/KajetanOzog/LLM_unbranding. The implementation of MUTE is available at https://github.com/KajetanOzog/MUTE.