SpatialOPSD:从经验证的编码智能体轨迹中自蒸馏空间智能
SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
SpatialOPSD是一种在线策略自蒸馏框架,通过将空间编码智能体的经验证轨迹作为特权信息,结合重复感知蒸馏技术,实现了无需外部工具的独立多模态大语言模型空间推理能力,在空间及分布外数据集上的性能优于SFT和GRPO。
AI中文摘要:
空间编码智能体通过使用外部工具生成经验证的执行轨迹,显著提升了多模态大语言模型(MLLM)的空间推理能力,但该范式固有存在推理时开销过高及依赖外部工具的问题。本文探究MLLM能否内化这种智能体能力以完全无需工具运行。我们首先发现,用空间编码智能体的摘要执行轨迹提示MLLM,可自然解锁模型内部的空间思维链(CoT)。受此启发,我们提出SpatialOPSD,这是一种在线策略自蒸馏框架,通过将经验证的智能体轨迹设定为特权信息,将空间推理内化到独立的MLLM中。为缓解蒸馏过程中的特权信息泄漏问题,我们引入重复感知蒸馏,其结合了重复掩码与非似然正则化。在多个基准上开展的实验表明,SpatialOPSD自蒸馏在空间数据集及分布外(OOD)数据集上的平均准确率均高于SFT和GRPO,展现出更优的性能与泛化能力。
英文摘要:
Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces. However, this paradigm inherently suffers from prohibitive inference-time overhead and external dependencies. In this paper, we explore whether an MLLM can internalize this agentic capability to operate entirely tool-free. We begin with a simple observation: prompting an MLLM with summarized execution traces of a spatial coding agent naturally unlocks the model's internal spatial Chain-of-Thought (CoT). Motivated by this, we introduce SpatialOPSD, an on-policy self-distillation framework that internalizes spatial reasoning into a standalone MLLM by formulating verified agent traces as privileged information. To mitigate privileged-information leakage during distillation, we introduce Repetition-Aware Distillation, which combines repetition masking with unlikelihood regularization. Experiments across multiple benchmarks demonstrate that self-distilling SpatialOPSD achieves higher average accuracy than SFT and GRPO on both spatial and OOD datasets, exhibiting superior performance and generalization.