AI 中文总结
该研究针对视觉-语言模型在细粒度空间任务上的不可靠问题,提出神经符号框架NeSy-Spatial,通过自演化空间技能提升工具利用精度与空间推理准确率。
AI 中文摘要
大型视觉-语言模型在多模态推理上已取得优异性能,但在需要精确空间感知和超出端到端生成的细粒度几何计算的细粒度空间任务上仍不可靠。工具增强是自然的解决方案,而现有方法要么从头规划工具调用且无显式依赖约束,要么依赖固定流程冗余且跨空间任务泛化性差。有效的空间推理智能体应积累可复用经验并自适应组合以解决新问题。为此,我们提出NeSy-Spatial,一种用于自演化空间技能的神经符号框架。NeSy-Spatial将工具交互和几何操作抽象为带类型的可执行原子指令,并将其组合为两种互补技能类型:用于组织工具执行的工具使用技能和用于结构化几何推理的几何技能。推理时,NeSy-Spatial在闭环过程中检索并执行相关技能;演化时,它分析缓冲的成功与失败轨迹以优化技能结构并修剪不可靠或不活跃的条目。在三个空间推理基准上的实验表明,NeSy-Spatial通过更精确的工具利用持续提升推理准确率。
英文摘要
Large vision-language models have achieved strong performance in multimodal reasoning, but they remain unreliable on fine-grained spatial tasks that demand both precise spatial perception and fine-grained geometric computation beyond end-to-end generation. Tool augmentation offers a natural solution, while existing methods either plan tool calls from scratch without explicit dependency constraints or rely on fixed pipelines that are redundant and generalize poorly across spatial tasks. An effective spatial reasoning agent should instead accumulate reusable experience and adaptively compose it for new problems. To this end, we propose NeSy-Spatial, a neuro-symbolic framework for self-evolving spatial skills. NeSy-Spatial abstracts tool interactions and geometric operations into typed executable atomic instructions and composes them into two complementary skill types: Tool-Use Skills for organizing tool execution and Geometry Skills for structured geometric reasoning. During inference, NeSy-Spatial retrieves and executes relevant skills in a closed-loop process. During evolution, it analyzes buffered successful and failed trajectories to refine skill structures and prune unreliable or inactive entries. Experiments on three spatial reasoning benchmarks show that NeSy-Spatial consistently improves reasoning accuracy with more precise tool utilization.
Comments25 pages, 9 figures, 10 tables; includes supplementary material