SpatialSkill:跨视角空间推理的自进化技能
SpatialSkill: Self-Evolving Skills for Cross-View Spatial Reasoning
浏览论文内容
中文总结 AI 辅助
SpatialSkill通过免权重更新框架,使冻结视觉-语言模型从离线轨迹积累显式自然语言空间推理技能,经视觉锚定与可执行性检查,在CityCube上跨四个执行器带来一致增益,9B模型超越最强闭源参考。
中文摘要 AI 辅助
跨视角空间推理要求模型将不同视角对齐为一致的空间表征,尽管这对人类而言是自然的,但这一能力对视觉-语言模型仍具挑战性。现有方法通常通过更新模型权重来改进空间推理,这使得所学知识保持隐式且与特定骨干网络绑定。我们提出SpatialSkill,一个免权重更新的框架,使冻结的视觉-语言模型能够从离线轨迹中积累显式的自然语言推理技能。与符号任务不同,感知技能无法仅通过执行来可靠验证:一个看似合理的空间规则可能缺乏视觉支持,或需要冻结模型无法执行的变换。因此,SpatialSkill仅在视觉锚定和可执行性检查后接纳候选技能,约束手动进化以防止有害回归,并按空间推理类别路由技能以减少负迁移。在CityCube上,跨四个冻结执行器,SpatialSkill带来一致的增益,配备SpatialSkill的9B执行器在我们评估中超过了最强的闭源参考。技能存储在带版本的自然语言手册中,使推理策略显式且可审计,无需修改模型参数。代码见https://this URL。
英文摘要
Cross-view spatial reasoning requires a model to align different viewpoints into a coherent spatial representation, yet this ability remains challenging for vision-language models despite being natural to humans. Existing methods typically improve spatial reasoning by updating model weights, which keeps the acquired knowledge implicit and tied to a specific backbone. We propose \textit{SpatialSkill}, a weight-update-free framework that enables a frozen vision-language model to accumulate explicit natural-language reasoning skills from offline trajectories. Unlike symbolic tasks, perceptual skills cannot be reliably verified simply by executing them: a plausible spatial rule may lack visual support or require transformations that the frozen model cannot perform. SpatialSkill therefore admits candidate skills only after visual-grounding and executability checks, constrains manual evolution to prevent harmful regressions, and routes skills by spatial-reasoning category to reduce negative transfer. On CityCube, across four frozen executors, SpatialSkill yields consistent gains, and a 9B executor equipped with SpatialSkill surpasses the strongest closed-source reference in our evaluation. The skills are stored in a versioned natural-language manual, making the reasoning strategies explicit and auditable without modifying model parameters. Code at https://github.com/vindahi/SpatialSkill.
发表机构
- Shandong University(山东大学)
- Shanghai AI Laboratory(上海人工智能实验室)
- Zhiyang Innovation Co., Ltd.(智洋创新科技股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。