视觉-语言模型智能体间潜在通信的事后稀疏编码
Post-Hoc Sparse Coding of Latent Communication Between Vision-Language Model Agents
浏览论文内容
中文总结 AI 辅助
该研究针对视觉-语言模型智能体间的潜在通信,通过事后稀疏自编码器压缩Vision Wormhole的传输数据,在保持性能的同时大幅减少了传输字节,为通信机制优化提供了方向。
中文摘要 AI 辅助
潜在空间通信使异构视觉-语言模型智能体能够交换连续表示,无需将视觉和推理状态序列化为文本。Vision Wormhole实现了这一方法,它将视觉特征转换为通用潜在表示,可供另一个模型使用,但每条消息都以相同大小的密集张量传输,无论其内容如何。因此,固定容量的密集张量不必具有固定的有效信息密度:某些消息可能仅使用了可用表示自由度的一小部分。这一观察结果表明,该通信信道可能具有相当大的可压缩性。我们通过对冻结的Vision Wormhole激活拟合事后稀疏自编码器(SAE),并在9个推理基准上测量重建、下游效用、特征复用和标记级干预来研究其冗余性。与原始float32传输相比,每个标记具有k=4个活跃系数的uint16索引/ float16值稀疏有效载荷将传输字节减少了128倍。在单次运行评估中,7项非AIME任务的平均准确率从49.85%变为49.77%。拟合的4096元素字典仅使用了50个特征,任务级活跃集的平均成对Jaccard相似度为0.906。这些测量结果确立了相对于原始传输的强大事后可压缩性,但尚未将稀疏编码的增量贡献与位置选择、精度降低、低秩结构或SAE优化效果分离开来。这些结果推动了匹配有效载荷比较以及其有效载荷适应每条消息所用信息的通信机制的发展。
英文摘要
Latent-space communication allows heterogeneous vision-language model agents to exchange continuous representations without serializing visual and reasoning states into text. Vision Wormhole realizes this approach by translating visual features into a universal latent representation that can be consumed by another model, but every message is transported as a dense tensor of the same size regardless of its content. A fixed-capacity dense tensor therefore need not have a fixed effective information density: some messages may use only a small fraction of the available representational degrees of freedom. This observation suggests that the communication channel may be substantially compressible. We study its redundancy by fitting a post-hoc sparse autoencoder to frozen Vision Wormhole activations and measuring reconstruction, downstream utility, feature reuse, and token-level interventions across nine reasoning benchmarks. Relative to the original float32 transport, a uint16-index/float16-value sparse payload with k=4 active coefficients per token reduces the transmitted bytes by 128x. In a single-run evaluation, the seven-task non-AIME mean accuracy changes from 49.85% to 49.77%. The fitted 4096-element dictionary uses only 50 features, and task-level active sets have a mean pairwise Jaccard similarity of 0.906. These measurements establish strong post-hoc compressibility relative to the original transport, but do not yet isolate the incremental contribution of sparse coding from position selection, reduced precision, low-rank structure, or SAE optimization effects. The results motivate matched-payload comparisons and communication mechanisms whose payload adapts to the information used by each message.
发表机构
- Xi’an Jiaotong-Liverpool University(西交利物浦大学)
机构由 AI 辅助整理,请以论文原文为准。