发表机构
IndigoWave; Sogang University(靛蓝波公司; 西江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出用于视频生成和机器人规划的4-张量注意力模型,在ROCStories数据集上测试,其性能优于基线模型且运行效率更高。
AI 中文摘要
我们提出一种4-张量注意力模型,该模型可预测场景的下一个语义状态,用于视频生成和机器人规划。状态窗口具有位置(x, t)以及两个纤维:语义纤维和时间上下文纤维,且采用一个softmax对窗口上的注意力进行联合归一化。帧和智能体的情境被表示为这些状态;编码器、渲染器和规划器保持在更新操作之外。为单独测试该更新操作,我们在ROCStories数据集上进行训练,其中每个窗口在语义层面临相同的下一句预测任务。在验证集上,每个设置使用一个随机种子,当H=2、L=2时,4-张量模型的最后一句交叉熵比自由运行的一维Transformer低5.3%;当H=4、L=2时,低2.6%;当H=4、L=3时,低2.4%。在H=4、L=2时,参数数量几乎相同,分别为1.725亿和1.759亿。在同两块GPU上,4-张量模型的运行时间为2.4小时,而基线模型的运行时间为45.2小时;基线模型通过自由运行解码进行训练,每个目标令牌对应一次顺序前向传播。
英文摘要
We describe a 4-tensor attention model that predicts the next semantic state of a scene, for video generation and robot planning. A window of states has positions (x, t) and two fibers, a semantic fiber and a temporal-context fiber, and one softmax normalizes attention jointly over the window. Frames and an agent's situation are written as those states; the encoder, the renderer, and the planner remain outside the update. To test the update on its own, we train on ROCStories, where each window poses the same next-sentence task at the semantic layer. On the validation split, with one seed per setting, the last-sentence cross-entropy on the three matched settings is lower for the 4-tensor model than for a free-running one-dimensional transformer by 5.3% at H=2, L=2, by 2.6% at H=4, L=2, and by 2.4% at H=4, L=3. At H=4, L=2 the parameter counts are nearly the same, 172.5M and 175.9M. On the same two GPUs that 4-tensor run finished in 2.4 hours and the baseline run in 45.2 hours; the baseline is trained by free-running decoding, one sequential forward pass per target token.
Comments36 pages, 4 figures