arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniAfford:面向可泛化2D-3D可供性感知的令牌路由多任务学习

UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception

Yuhao Liu, Yiming Zhong, Hanqing Wang, Shaocheng Yan, Yuhang Zhang, Wenzhou Lyu, Ziyang Ding, Wei Zhang, Xue Zhao, Jin Pan, Yuexin Ma, Xinge Zhu

arXiv 2609.37264首次发表:更新:

发表机构

ShanghaiTech University; Shandong University; Yinwang Intelligent Technology Co., Ltd.; HKUST(GZ); CUHK; Wuhan University(上海科技大学; 山东大学; 银网智能科技有限公司; 香港科技大学(广州); 香港中文大学; 武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出UniAfford框架,通过任务令牌路由器实现MLLM的多任务学习,统一2D-3D可供性感知,支持异构监督和零样本泛化,在基准上达到最先进性能。

AI 中文摘要

可供性感知旨在定位支持具身交互的可操作区域,然而2D和3D可供性定位已演变为独立的问题,具有不同的任务定义、监督格式、数据集和评估协议。这种碎片化限制了跨视觉和几何空间的可迁移对象-可供性语义的学习。我们提出了任务令牌路由器,一种基于MLLM系统的多任务训练范式,它将上下文隐藏状态路由到任务特定分支,而无需语言头生成预定义标记。路由状态直接由分支特定目标监督,使密集预测损失能够塑造共享的MLLM表示。我们将此范式实例化为UniAfford,一个用于可泛化2D-3D可供性感知的统一框架,以及UniAfford-Data,一个在共享对象-可供性分类法下整合像素级2D注释、点级3D注释和语言指令的数据集,通过语义级2D-3D配对支持异构监督。UniAfford采用MLLM作为共享语义中心,并使用模态感知令牌路由器生成图像和点云可供性查询。这些查询分别条件化SAM风格的像素解码器和基于SONATA的点解码器,实现从仅图像、仅点云或配对多模态输入的灵活2D、3D和联合可供性推理。实验表明,在2D和3D可供性基准上,无需目标特定微调即可实现强大的零样本泛化,同时在模态隔离协议下取得最先进的分支性能。消融研究验证了令牌路由、联合2D-3D监督和解码器耦合,而语言头诊断显示路由的潜在状态携带有意义的对象-可供性语义。项目页面:此https URL

英文摘要

Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separate problems, with different task definitions, supervision formats, datasets, and evaluation protocols. This fragmentation limits the learning of transferable object-affordance semantics across visual and geometric spaces. We propose Token Router for Tasks, a multitask training paradigm for MLLM-based systems that routes contextual hidden states to task-specific branches without requiring the language head to generate predefined markers. Routed states are supervised directly by branch-specific objectives, enabling dense prediction losses to shape shared MLLM representations. We instantiate this paradigm as UniAfford, a unified framework for generalizable 2D-3D affordance perception, together with UniAfford-Data, a dataset integrating pixel-level 2D annotations, point-level 3D annotations, and language instructions under a shared object-affordance taxonomy, supporting heterogeneous supervision through semantic-level 2D-3D pairing. UniAfford adopts an MLLM as a shared semantic hub and a modality-aware token router to produce image- and point-cloud-affordance queries. These queries respectively condition a SAM-style pixel decoder and a SONATA-based point decoder, enabling flexible 2D, 3D, and joint affordance inference from image-only, point-cloud-only, or paired multimodal inputs. Experiments demonstrate strong zero-shot generalization across 2D and 3D affordance benchmarks without target-specific fine-tuning, alongside state-of-the-art branch-wise performance under modality-isolated protocols. Ablations validate token routing, joint 2D-3D supervision, and decoder coupling, while language-head diagnostics show that routed latent states carry meaningful object-affordance semantics. Project page: https://4dvlab.github.io/UniAfford

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑