InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
Nianchen Deng, Lixin Gu, Shenglong Ye, Yinan He, Zhe Chen, Songze Li, Haomin Wang, Xingguang Wei, Tianshuo Yang, Min Dou, Tong He, Wenqi Shao, Kaipeng Zhang, Yi Wang, Botian Shi, Yanting Zhang, Jifeng Dai, Yu Qiao, Hongjie Zhang, Wenhai Wang
机构
*
Shanghai AI Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong(香港中文大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Donghua University(东华大学)
;
Nanjing University(南京大学)
;
Tsinghua University(清华大学)
RationalVLA: A Rational Vision-Language-Action Model with Dual System
Wenxuan Song, Jiayi Chen, Wenxue Li, Xu He, Han Zhao, Can Cui, Pengxiang Ding Shiyan Su, Feilong Tang, Xuelian Cheng, Donglin Wang, Zongyuan Ge, Xinhu Zheng, Zhe Liu, Hesheng Wang, Haoang Li
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Westlake University(西湖大学)
;
Monash University(墨尔本大学)
;
Shanghai Jiao Tong University(上海交通大学)
CHIP: A multi-sensor dataset for 6D pose estimation of chairs in industrial settings
Mattia Nardon, Mikel Mujika Agirre, Ander González Tomé, Daniel Sedano Algarabel, Josep Rueda Collell, Ana Paola Caro, Andrea Caraffa, Fabio Poiesi, Paul Ian Chippendale, Davide Boscaini
QueryCAD: Grounded Question Answering for CAD Models
Claudius Kienle, Benjamin Alt, Darko Katic, Rainer Jäkel, Jan Peters
机构
*
ArtiMinds Robotics
;
IAS Lab, Computer Science Department, TU Darmstadt(IAS实验室,计算机科学系,图腾斯泰特大学)
;
AICOR Institute for Artificial Intelligence, University of Bremen(AICOR人工智能研究所,不莱梅大学)
;
Stuttgart University of Applied Sciences(斯图加特应用科学大学)
CommentsAccepted at 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); Fifth International Workshop on Event-Based Vision
Journal refIEEE Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, 2025
Comments7 pages, 4 figures. Submitted to the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2025. This work has been submitted to the IEEE for possible publication
Detecting socially interacting groups using f-formation: A survey of taxonomy, methods, datasets, applications, challenges, and future research directions
CommentsKun Jiang, Mengmeng Yang and Diange Yang are Corresponding Author. The main paper and supplementary material are both included here, total 23 pages (main paper is 10 pages and supplementary material is 13 pages), total 17 figures (6 figures in main paper and 11 figures in supplementary material), this paper is Accepted to CVPR WDFM-AD Workshop 2025, The code will be available at https://Ankit-Zefan.github.io/CleanMap/