arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MORPH:基于神经可塑性启发的自适应拓扑的多机器人自组织任务分配

MORPH: Self-Organising Multi-Robot Task Allocation via Neuroplasticity-Inspired Adaptive Topology

Xuezhi Niu, Didem Gürdür Broo

arXiv 2609.32745首次发表:更新:

AI 中文总结

MORPH提出无需训练的MRTA框架,通过四条局部可塑性规则在线学习有向偏好矩阵,在TA-RWARE基准上以21%的协调链接实现110%吞吐量,且对任务分布偏移鲁棒性优于邻近度方法。

AI 中文摘要

动态环境中的多机器人任务分配(MRTA)面临一个基本矛盾:有效的协调需要学习到的结构,但这种结构必须在条件变化时适应。现有方法通过假设先验任务知识、效用函数、成本矩阵或训练好的策略来解决这一问题,这使得它们在缺乏此类知识部署或任务分布发生变化时变得脆弱。我们提出MORPH(通过可塑性引导的层次结构进行多智能体在线重连),一种无需训练的MRTA框架,其中全局分配质量由4条局部可塑性规则(突触、稳态、结构和元可塑性)涌现,这些规则应用于从运行时共现和任务完成反馈更新的有向成对偏好矩阵。MORPH不需要任务模型、不需要出价计算、也不需要离线训练;响应决策使用学习到的AGV到拣选员的偏好,而非固定的邻近规则。在Gerkey-Mataric MRTA分类中,MORPH是单任务、单机器人、即时分配类别中首个在线学习有向成对分配偏好的方法。在TA-RWARE仓库基准(8-24个智能体、4张地图、每回合800步、5个种子)上评估,MORPH在N=24时达到全对全吞吐量的110%,同时仅使用21%的可能协调链接,这一效率优势随车队规模单调增长。在空间任务分布偏移下,MORPH的退化程度比基于邻近度的方法低3倍,而其学习到的偏好与曼哈顿距离保持不相关。系统性消融实验证实所有四条可塑性规则均有可测量的贡献。两种分配特性无需编程即涌现:跨类型偏好主导和渐进式偏好稀疏化,反映了生物神经回路的发育性细化。学习到的偏好由任务共现历史驱动,而非空间邻近度。

英文摘要

Multi-robot task allocation (MRTA) in dynamic environments faces a fundamental tension: effective coordination requires learned structure, but that structure must adapt when conditions change. Existing methods resolve this by assuming prior task knowledge, a utility function, a cost matrix, or a trained policy making them brittle when deployed without such knowledge or when task distributions shift. We present MORPH(Multi-agent Online Rewiring through Plasticity-guided Hierarchy), a training-free MRTA framework where global allocation quality emerges from 4 local plasticity rules (synaptic, homeostatic, structural, and metaplasticity) applied to a directed pairwise preference matrix updated from runtime co-occurrence and task-completion feedback. MORPH requires no task model, no bid computation, and no offline training; response decisions use learned AGV-to-Picker preferences rather than a fixed proximity rule. Within the Gerkey-Mataric MRTA taxonomy, MORPH is the first method in the single-task, single-robot, instantaneous-assignment class to learn directed pairwise allocation preferences online. Evaluated on the TA-RWARE warehouse benchmark (8-24 agents, 4 maps, 800 steps per episode, 5 seeds), MORPH achieves 110% of all-to-all throughput at N=24 while using only 21% of possible coordination links as an efficiency advantage that grows monotonically with fleet size. Under spatial task distribution shift, MORPH degrades 3x less than proximity-based methods while its learned preferences remain uncorrelated with Manhattan distance. Systematic ablation confirms all four plasticity rules contribute measurably. Two allocation properties emerge without programming: cross-type preference dominance and progressive preference sparsification, mirroring the developmental refinement of biological neural circuits. Learned preferences are driven by task co-occurrence history, not spatial proximity.

CommentsFull version of the paper published in The 23rd European Conference on Multi-Agent Systems (EUMAS), 2026.The implementation and experiment scripts are available at https://github.com/Cyber-physical-Systems-Lab/morph_v2

Journal refIn Multi-Agent Systems, EUMAS 2026, LNCS 16829, Springer, 2026

DOI:10.1007/978-3-032-39395-1_31

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑