arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

单未来,每机器人:采用去中心化JEPA的标签高效集体状态预测

One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA

Alan-Barsag Gazzaev, Alexey Gavrilov, Sergey Muravyov

arXiv 2607.28443首次发表:更新:

发表机构

ITMO University(ITMO大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出CS-JEPA架构,解决群体机器人去中心化共享状态预测问题,在多类场景下提升预测误差、一致性及相关指标,为去中心化群体预测提供标签高效基元。

AI 中文摘要

群体中的每个机器人能否仅通过局部观测和带宽受限的消息,预测出相同的未来集体状态?我们将此问题表述为去中心化共享状态预测,并提出集体状态联合嵌入预测架构(Collective-State JEPA,CS-JEPA),这是一种循环联合嵌入预测架构,每个机器人的输出代表一个共同的未来令牌域。部署时,每个机器人使用16帧局部历史和每条有向边的一个64浮点循环消息;不存在全局池化、目标编码器、回合时钟或记录的未来动作。在未使用下游集体标签进行预训练后,冻结的表示通过在6、12或24个全局标记回合上拟合的岭探针进行评估。与采用相同接收者锚点和部署容量但额外增加9607个仅训练参数的原始未来重建方法相比,一项前瞻性注册的五种子后续研究在分布内、环形、互k近邻及高达108个机器人的未见过规模族上,将预测误差和机器人间一致性标签预算AUC提升了最多108个百分点。5/5个外部种子的所有效果均有利于CS-JEPA。在另一项密封的八种子后续研究中,匹配的动作条件预测器在生成接收者局部预测表示前,会接收每个候选的四步计划。CS-JEPA将分支值均方误差(MSE)降低45.5%,并将上下文内候选得分皮尔逊相关系数提高0.1291,8/8个种子的两种效果均有利,包括在未见过的N=32时。这些结果表明,共同未来JEPA目标是拓扑和规模变化下去中心化群体预测的一种标签高效基元,同时为与规划相关的价值估计提供了额外证据。

英文摘要

Decentralized robots often need a common view of what their team is becoming, even though each robot sees different evidence and cannot rely on a central estimate or output-level consensus. We ask whether compatible collective-state predictions can emerge under this constraint. Collective-State JEPA (CS-JEPA) trains every robot to predict the same fixed-width latent future from its own history and bounded neighbor messages, with no agreement loss; predictions and plans are never pooled at deployment. In a fresh independent replication, agreement improves for every seed and every evaluated split. Accuracy improves at the same time, ruling out the uninformative solution in which all robots merely collapse to one prediction: relative to capacity-matched raw-future reconstruction, collective-state error falls by 28.4 percent in distribution and by 64.4 to 75.6 percent under topology and swarm-size shift. Translation-free and crossed-pretraining controls preserve this joint result, while action-conditioned and rigid-body evaluations show that the receiver-local representation supports independent decisions. A shared latent future can therefore align decentralized predictions without consensus training while preserving useful, label-efficient information.

CommentsSubmitted to IEEE ICRA 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑