arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SpikeWorld:用于冻结脉冲世界模型的快速状态适应

SpikeWorld: Fast-State Adaptation for Frozen Spiking World Models

Ziqiao Yu

arXiv 2608.07712首次发表:更新:

发表机构

DiDi International Business Group(滴滴国际业务集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SpikeWorld是145万参数的稀疏脉冲模型,冻结训练参数后通过外部路径实现快速状态适应,提升多模态预测等性能,在Meta-World任务中提高冻结策略奖励与成功率。

AI 中文摘要

每当观察到动作的结果时,预测模型会接收到自监督信号。当动力学和语义共享参数时,部署后使用该信号很困难:冻结会阻止适应,而权重更新需要优化器状态,可能会改变学习到的表示。在此,我们引入SpikeWorld,这是一个拥有145万参数的稀疏脉冲模型,经过联合训练,用于异构感官预测、语义、图像-文本绑定和动作条件动力学。在部署时,所有训练好的参数均被冻结。延迟的下一个状态残差更新两条外部路径:累积固定库损失选择有界动作校正,而特定路径的残差矩阵则细化下一个状态预测。两条路径均不使用标签、教师输出、奖励、成功信号或真实偏移值。联合优化使动作下一个状态的MSE降低17.10%,同时还改善了多模态预测、语义准确率和图像-文本检索。在保留的剪切和衰减流上,组合的外部状态使总预测分别提高5.48%和30.01%;其固定库动作路径的跟踪性能分别提高24.20%和3.94%。在包含450条新Meta-World轨迹(每条臂75条)的六臂研究中,SpikeWorld使冻结策略的奖励提高了7.90(95%置信区间[2.48, 14.06]);13.33点的成功率差异具有描述性(置信区间[0, 40])。对于相同的感官输入,模型参数和继承的语义输出保持逐位不变。一个16字节的RLS估计器在线性衰减上获得了最高的非oracle奖励,表明该贡献并非优越的线性识别,而是其与冻结的多模态脉冲检查点的集成。参考代码可在该https URL公开获取。

英文摘要

A predictive model receives a self-supervised signal whenever the consequence of an action is observed. Using that signal after deployment is difficult when dynamics and semantics share parameters: freezing prevents adaptation, whereas weight updates require optimizer state and may alter the learned representation. Here we introduce SpikeWorld, a 1.45M-parameter sparse spiking model jointly trained for heterogeneous sensory prediction, semantics, image-text binding and action-conditioned dynamics. At deployment, all trained parameters are frozen. Delayed next-state residuals update two external paths: cumulative fixed-bank losses select the bounded action correction, while route-specific residual matrices refine next-state prediction. Neither path uses labels, teacher outputs, rewards, success signals or the true shift value. Joint optimization improves action next-state MSE by 17.10\% while also improving multimodal prediction, semantic accuracy and image-text retrieval. On held-out shear and attenuation streams, the combined external state improves aggregate prediction by 5.48\% and 30.01\%; its fixed-bank action path improves tracking by 24.20\% and 3.94\%, respectively. In a six-arm study comprising 450 new Meta-World trajectories (75 per arm), SpikeWorld raises frozen-policy reward by 7.90 (95\% CI [2.48, 14.06]); the 13.33-point success difference is descriptive (CI [0, 40]). For identical sensory inputs, model parameters and inherited semantic outputs remain bitwise unchanged. A 16-byte RLS estimator obtains the highest non-oracle reward on linear attenuation, showing that the contribution is not superior linear identification, but its integration with a frozen multimodal spiking checkpoint. Reference code is publicly available at https://github.com/Oooorca/SpikeWorld.

Comments14 pages, 2 figures, 4 tables. Code: https://github.com/Oooorca/SpikeWorld

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑