HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV
视觉与机器人
三维重建、NeRF、Gaussian Splatting、点云和空间智能。
专题命中 空间理解 :spatial understanding(abstract);分类 cs.CV
机构 * Wangxuan Institute of Computer Technology, Peking University(计算机技术王先院,北京大学) ; National Institute of Health Data Science, Peking University(健康数据科学国家研究院,北京大学) ; State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)
专题命中 空间理解 :3D vision(abstract);分类 cs.CV
Comments Accepeted to ACM MM 25