机构
*
Peking University(北京大学)
;
Northeastern University(东北大学)
;
Sydney Smart Technology College(悉尼智能技术学院)
;
School of Computer Science, Peking University(北京大学计算机学院)
;
The State Key Laboratory of Multimedia Information Processing(多媒体信息处理国家重点实验室)
Grasping Deformable Objects via Reinforcement Learning with Cross-Modal Attention to Visuo-Tactile Inputs
Yonghyun Lee, Sungeun Hong, Min-gu Kim, Gyeonghwan Kim, Changjoo Nam
机构
*
Dept. of Electronic Engineering at Sogang University(ソガン大学电子工程系)
;
Dept. of Immersive Media and Engineering at Sungkyunkwan University(顺天大学沉浸媒体与工程系)
;
College of Medicine, Yonsei University(延世大学医学院)
Multimodal Disease Progression Modeling via Spatiotemporal Disentanglement and Multiscale Alignment
Chen Liu, Wenfang Yao, Kejing Yin, William K. Cheung, Jing Qin
机构
*
School of Nursing, The Hong Kong Polytechnic University(香港理工大学护理学院)
;
Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学系)
Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model
Xinyue Lou, You Li, Jinan Xu, Xiangyu Shi, Chi Chen, Kaiyu Huang
机构
*
Key Laboratory of Big Data & Artificial Intelligence in Transportation (Beijing Jiaotong University), Ministry of Education(大数据与人工智能交通联合实验室(北京交通大学))
;
School of Computer Science and Technology, Beijing Jiaotong University(计算机科学与技术学院,北京交通大学)
;
Tsinghua University(清华大学)
机构
*
EPIC Lab, Shanghai Jiao Tong University(上海交通大学EPIC实验室)
;
Sichuan University(四川大学)
;
University of Electronic Science & Technology of China(电子科技大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Sun Yat-sen University(中山大学)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
NV3D: Leveraging Spatial Shape Through Normal Vector-based 3D Object Detection
Krittin Chaowakarn, Paramin Sangwongngam, Nang Htet Htet Aung, Chalie Charoenlarpnopparut
机构
*
The School of Information, Computer, and Communication Technology, Sirindhorn International Institute of Technology, Thammasat University(信息、计算机与通信技术学院,Sirindhorn国际技术学院,泰国朱拉隆梭大学)
;
National Electronics and Computer Technology Center, National Science and Technology Development Agency(国家电子与计算机技术中心,国家科学技术发展局)
;
Department of Electrical Engineering, Faculty of Engineering, Chulalongkorn University(电气工程系,工程学院,朱拉隆梭大学)
Isolated Channel Vision Transformers: From Single-Channel Pretraining to Multi-Channel Finetuning
Wenyi Lian, Patrick Micke, Joakim Lindblad, Nataša Sladoje
机构
*
Department of Information Technology Uppsala University(信息科技系乌普萨拉大学)
;
Department of Immunology, Genetics and Pathology Uppsala University(免疫学、遗传学和病理学系乌普萨拉大学)
专题命中
多模态训练与对齐
:multimodal(abstract);分类 cs.CV
CommentsPaper has been accepted by BMVC as an Oral presentation
机构
*
Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)
;
School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
;
CTTL-Terminal, China Academy of Information and Communications Technology(信息通信技术中国科学院CTTL终端)
;
School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院)
专题命中
多模态训练与对齐
:multi-modal(abstract);分类 cs.CV
Comments10 pages, 6 figures, Accepted by ACM MM 2025
CommentsAccepted and Published in SBP-BRiMS 2025. 18th International Conference on Social Computing, Behavioral-Cultural Modeling & Prediction and Behavior Representation in Modeling and Simulation