arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00502cs.CV

SpatialAfford:教紧凑视觉语言模型(VLMs)关注何处与定位何处以实现 affordance

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance

  • Zhejiang University(浙江大学)
  • Vivo Mobile Communication Co., Ltd(维沃移动通信有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Yufei Zhang, Chenlu Zhan, Donghui Sun, Xiaoxin Chen, Hongwei Wang

AI总结:

SpatialAfford 是一个两阶段框架,通过空间注意力对齐和空间感知 GRPO,使紧凑 4B VLMs 在多个基准上的 affordance 定位性能优于更强的 7B+ 基线模型。

AI中文摘要:

affordance 定位旨在定位用于交互的功能区域,例如用于抓取的把手或用于按压的按钮,而非整个物体。这比通用视觉定位更具挑战性,因为目标区域更小、更模糊且更依赖任务上下文,尤其适用于具身环境中使用的紧凑视觉语言模型(VLMs)。近期的序列级监督和强化学习提升了坐标预测质量,但紧凑自回归 VLMs 在坐标生成前仍缺乏可靠的 affordance 感知视觉聚焦:模型可生成更好的坐标 token,但其跨模态注意力仍呈弥散状且未牢固锚定到真实 affordance 证据。为解决该问题,我们提出 SpatialAfford,这是一个两阶段框架,首先通过空间注意力对齐(Spatial Attention Alignment, SAA)将注意力对齐到真实 affordance 区域,随后通过空间感知 GRPO 优化坐标预测。通过在优化定位何处之前明确教导模型关注何处,SpatialAfford 将 affordance 定位从纯粹的输出约束目标转变为基于注意力的空间推理。在 ShareRobot-Bench、ReasonAff 和 PartAfford 数据集上,SpatialAfford 持续提升 affordance 定位性能,其中一个紧凑的 4B 模型的表现优于更强的 7B+ 基线模型。

英文摘要:

Affordance grounding aims to localize the functional region for interaction, such as the handle to grasp or the button to press, rather than the whole object. This makes it more challenging than generic visual grounding because the target region is smaller, more ambiguous, and more dependent on task context, especially for compact vision-language models (VLMs) used in embodied settings. Recent sequence-level supervision and reinforcement learning improve coordinate prediction quality, yet compact autoregressive VLMs still lack reliable affordance-aware visual focus before coordinate generation: the model can produce better coordinate tokens while its cross-modal attention remains diffuse and weakly anchored to the true affordance evidence. To address it, we propose SpatialAfford, a two-stage framework that first aligns attention to the ground-truth affordance region through Spatial Attention Alignment (SAA), then refines coordinate prediction with Spatial-Aware GRPO. By explicitly teaching the model where to look before optimizing where to ground, SpatialAfford turns affordance grounding from a purely output-constrained objective into attention-grounded spatial reasoning. Across ShareRobot-Bench, ReasonAff, and PartAfford, SpatialAfford consistently improves affordance grounding, with a compact 4B model outperforming stronger 7B+ baselines.

↑