发表机构
EPFL; University of California, San Diego(洛桑联邦理工学院; 加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对去中心化LLM微调中的传播后门攻击,提出Chorus机制,通过邻居独立探测与投票检测并拒绝恶意适配器,将平均攻击成功率从48-63%降至2.2%以内,且通信开销可忽略。
AI 中文摘要
去中心化大语言模型(LLM)微调允许组织在无法汇集数据的情况下,无需中央协调者即可协作训练共享的LLM。在每一轮中,每个节点通过通信图与其邻居交换可训练适配器,然后进行聚合。然而,这种设置容易受到传播后门的影响,后门是一种隐藏行为,使模型在干净输入上表现正常,但每当出现秘密触发器时便产生攻击者选择的输出。我们证明,单个节点毒化其自身模型即可使从未见过中毒示例的节点的适配器被植入后门,使其拒绝包含秘密触发器的提示。我们提出了Chorus,一种去中心化机制,使每个节点在聚合前能够检测并拒绝来自其邻居的后门适配器,无需共享验证数据或了解攻击者的触发器或目标。Chorus根据每个适配器的行为进行评判,以接收者自身的适配器作为可信参考。关键在于,Chorus中没有节点单独评判适配器:每个适配器的接收者独立探测它,在邻域内汇总其发现,并投票做出决策。因此,一个逃过一个接收者的后门仍会被其他接收者捕获。我们使用两个指令微调数据集和LLM架构,并针对最先进的基线评估了Chorus的有效性。Chorus将攻击者邻居的平均攻击成功率(ASR)从48-63%降至至多2.2%,与知晓确切恶意节点的全知预言机相差0.6个百分点以内。即使受影响最严重的诚实节点,其ASR也从未超过10%,与预言机相同,而相比之下无防御时高达78%。这一切仅带来可忽略的通信开销。
英文摘要
Decentralized large language model (LLM) fine-tuning lets organizations collaboratively train a shared LLM on data they cannot pool, without a central coordinator. In every round, each node exchanges a trainable adapter with its neighbors over a communication graph, and then aggregates them. This setting, however, is vulnerable to propagated backdoors, which is a hidden behavior that lets a model perform normally on clean inputs but produce an attacker-chosen output whenever a secret trigger appears. We show that a single node poisoning its own model can backdoor adapters of nodes that have never seen a poisoned example, making them refuse prompts that contain a secret trigger. We present Chorus, a decentralized mechanism that lets each node detect and reject backdoored adapters from its neighbors before aggregation, without requiring shared validation data or any knowledge of the attacker's trigger or target. Chorus judges each adapter by its behavior, using the receiver's own adapter as a trusted reference. Crucially, no node in Chorus judges adapters alone: the receivers of each adapter update probe it independently, pool their findings in the neighborhood, and vote to make a decision. So a backdoor that slips past one receiver is still caught by the others. We evaluate the effectiveness of Chorus using two instruction-tuning datasets and LLM architectures, and against a state-of-the-art baseline. Chorus cuts the average attack success rate (ASR) of the attacker's neighbors from 48-63% to at most 2.2%, within 0.6 percentage points of an omniscient oracle that knows the exact malicious nodes. Even the worst-affected honest node never exceeds 10% ASR, the same bound as the oracle, against up to 78% without defense. This all comes at a negligible communication overhead.