arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36770cs.AIcs.CVcs.LG

无共享门控或跨智能体梯度的自监督协作视觉专家群体中的涌现特化

Emergent Specialization in Populations of Self-Supervised Collaborative Vision Experts Without a Shared Gate or Cross-Agent Gradients

Aram Davtyan, Pablo Acuaviva, Sebastian Stapf, Paolo Favaro

首次发表
浏览论文内容

中文总结 AI 辅助

该研究证明,在没有共享门控或跨智能体梯度的条件下,自监督协作视觉专家群体可通过局部路由和无梯度通信涌现出有用的特化分工,提升性能并超越单一通才模型。

中文摘要 AI 辅助

一个神经网络群体能否在没有共享门控或智能体间梯度的情况下发展出有用的分工?我们研究了一个设置,其中每个网络拥有自己的权重,在相同的异构数据上独立训练,并可以通过前向传播向另一个智能体请求帮助。与混合专家模型(其中联合训练的门控将输入分配给专家)不同,这里的特化必须在没有中央控制的情况下涌现。我们在预测性视觉预训练的小规模代理中对此进行了测试。最初相同的智能体使用冻结的DINOv3特征的掩码预测,在六个视觉域的无标签混合数据上进行微调。我们通过询问某个输入的最佳智能体是否与其潜在域对齐来衡量特化,并通过询问责任是否在智能体间分布来衡量利用率。我们逐步移除中央控制,最终得到DISCO(分布式协作),其中每个智能体局部选择助手,通过无梯度通道读取其内部状态,并仅根据帮助带来的改进对其路由器进行奖励。特化涌现且有用。随机路由的群体性能低于单一通才,而语义路由的群体性能优于通才,这表明特化而非群体规模驱动了性能提升。特化在没有中央路由器的情况下持续存在,无梯度通信使非专家能够利用涌现的专业知识。在DISCO中,由专家帮助的随机智能体与单独的通才表现相当,而专家则超越通才,包括在特化混合之外的数据上。局部路由器为98%的输入选择涌现的专家。这些效应在群体规模、模型容量、数据不平衡和微调种子下持续存在,为去中心化预测性预训练所需的动态提供了可测量的证据。

英文摘要

Can a population of neural networks develop a useful division of labor without a shared gate or gradients between agents? We study a setting where each network has its own weights, trains independently on the same heterogeneous data, and can ask another agent for help through a forward pass. Unlike mixtures of experts, where a jointly trained gate assigns inputs to experts, specialization here must emerge without central control. We test this in a small scale proxy for predictive visual pretraining. Initially identical agents are finetuned on an unlabeled mixture of six visual domains using masked prediction of frozen DINOv3 features. We measure specialization by asking whether the best agent for an input aligns with its latent domain, and utilization by asking whether responsibility is distributed across agents. We progressively remove central control, ending with DISCO (DIStributed COllaboration) where each agent locally selects a helper, reads its internal state through a gradient free channel, and rewards its router only for the improvement that help provides. Specialization emerges and is useful. Randomly routed populations underperform a single generalist, while semantically routed populations outperform it, showing that specialization rather than population size drives the gain. Specialization persists without a central router, and gradient free communication lets nonexperts exploit emergent expertise. In DISCO, a random agent helped by the expert matches the solo generalist, while experts surpass it, including on data outside the specialization mixture. Local routers select the emergent expert for 98% of inputs. These effects persist across population size, model capacity, data imbalance, and finetuning seeds, providing measurable evidence for the dynamics needed by decentralized predictive pretraining.

发表机构

  • University of Bern(伯尔尼大学)

机构由 AI 辅助整理,请以论文原文为准。

↑