XBridge:面向异构大语言模型通信的实体锚定隐空间桥接器
XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication
- Hankuk University of Foreign Studies(韩国外国语大学)
- Noah’s Farm(诺亚农场)
- University of Illinois Chicago(伊利诺伊大学芝加哥分校)
- Capital One AI Foundations(Capital One AI 基金会)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
XBridge是免解码的异构大语言模型通信协议,通过词汇锚定映射和隐空间丰富桥接器解决跨架构通信的实体身份丢失问题,在多模型、多基准测试中表现优于现有方法,参数开销低。
AI中文摘要:
由不同模型家族驱动的异构多智能体大语言模型系统,可通过减少冗余推理模式,表现优于同构配置。然而现有通信协议要么通过文本传输,丢弃发送方的内部表征;要么要求架构同构才能进行隐空间层面的传递。我们在跨架构通信中发现了实体锚定问题:跨注意力桥接器在不同大语言模型家族间传递连续表征时,会出现稀有令牌压缩崩溃,实体身份在连续瓶颈(仅桥接器的F1值约为30%)中丢失。我们提出XBRIDGE,一种免解码的通信协议,通过两种机制解决该问题:词汇锚定映射(Lexical Anchor Mapping, LAM)将发送方的原始上下文令牌映射到接收方的词汇表,提供离散实体锚点;隐空间丰富桥接器(Latent Enrichment Bridge, LEB)让接收方查询发送方的隐藏状态以丰富上下文。实体锚点通过接收方自身的自注意力,将桥接器的上下文信号锚定到特定实体。在三个模型家族(Llama、Qwen和Mistral)、七个基准测试及双向通信场景中,XBRIDGE在每对模型的所有七个任务上均优于基于文本的通信,同时延迟降低11倍;在同架构设置下,其在七个任务中的六个上也优于键值共享基线。LEB仅需2.64亿个可训练参数(为接收方参数的3.8%),在小型平衡样本集上训练,推理开销可忽略不计。
英文摘要:
Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by reducing redundant reasoning patterns. Yet existing communication protocols either operate through text, discarding the sender's internal representations, or require architectural homogeneity for latent-level transfer. We identify the entity grounding problem in cross-architecture communication: cross-attention bridges that transfer continuous representations across different LLM families suffer from rare-token compression collapse, where entity identity is lost in the continuous bottleneck (bridge-only F1 ~30%). We propose XBRIDGE, a decode-free communication protocol that addresses this through two mechanisms. Lexical Anchor Mapping (LAM) maps the sender's original context tokens to the receiver's vocabulary, providing discrete entity anchors. A Latent Enrichment Bridge (LEB) lets the receiver query the sender's hidden states for contextual enrichment. The entity anchors ground the bridge's contextual signals to specific entities through the receiver's own self-attention. Across three model families (Llama, Qwen, and Mistral), seven benchmarks, and both communication directions, XBRIDGE outperforms text-based communication on all seven tasks for each model pair while achieving 11x lower latency, and in a same-architecture setting it also exceeds a KV-sharing baseline on six of seven tasks. LEB requires only 264M trainable parameters (3.8% of the receiver), is trained on a small balanced sample set, and adds negligible inference overhead.