机构
*
University of Technology Sydney(悉尼科技大学)
;
Nanjing University of Science and Technology(南京理工大学)
;
Honor Device Co., Ltd.(荣耀终端有限公司)
;
Shanghai Institute of Microsystem and Information Technology, CAS(中国科学院上海微系统与信息技术研究所)
机构
*
Institute of Computing Technology, CAS(中国科学院计算技术研究所)
;
Hangzhou Dianzi University(杭州电子科技大学)
;
North China Electric Power University(华北电力大学)
;
Lishui Institute of Hangzhou Dianzi University(杭州电子科技大学丽水学院)
MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding
MAC 2026:推动微动作分析迈向细粒度理解
Kun Li, Dan Guo, Jihao Gu, Pengyu Liu, Xiaobai Li, Haoyu Chen, Yanbin Hao, Guoying Zhao, Meng Wang
机构
*
United Arab Emirates University(阿联酋大学)
;
Hefei University of Technology(合肥工业大学)
;
University College London(伦敦大学学院)
;
Zhejiang University(浙江大学)
;
University of Oulu(奥卢大学)
;
CMVS, University of Oulu(奥卢大学计算机视觉与媒体研究中心)
CommentsAuthor Accepted Manuscript. Accepted for publication in the Proceedings of the 34th ACM International Conference on Multimedia (ACM MM '26). This author-created manuscript is not the ACM Version of Record
机构
*
School of Computer Science, Peking University(北京大学计算机科学学院)
;
State Key Laboratory for Multimedia Information Processing, Peking University(北京大学多媒体信息处理国家重点实验室)
;
Academy for Advanced Interdisciplinary Studies, Peking University(北京大学前沿交叉学科研究院)
;
Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)
MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
MMaDA-VLA: 基于统一多模态指令与生成的大型扩散视觉-语言-动作模型
Yang Liu, Pengxiang Ding, Tengyue Jiang, Xudong Wang, Wenxuan Song, Minghui Lin, Han Zhao, Hongyin Zhang, Zifeng Zhuang, Wei Zhao, Siteng Huang, Jinkui Shi, Donglin Wang
机构
*
Westlake University(西湖大学)
;
Zhejiang University(浙江大学)
;
East China University of Science and Technology(华东理工大学)
;
Huawei Celia Team(华为Celia团队)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
OpenHelix Robotics
CommentsAccepted to appear in the Proceedings of the 34th ACM International Conference on Multimedia (MM '26). 10 pages, 9 figures, and 2 tables. Code: https://github.com/semcomm/GVC-RT
机构
*
The University of Tokyo(东京大学)
;
Nara Institute of Science and Technology(奈良科学技术研究所)
;
Chungnam National University(忠南国立大学)
;
Institute of Science Tokyo(东京科学大学)
ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination
ReGround:通过自诊断与视觉重检验恢复多步推理中的视觉接地
Lei Peng, Shuai Lv, Wei Hu
机构
*
University of Science and Technology of China(中国科学技术大学)
;
School of Artificial Intelligence and Data Science(人工智能与数据科学学院)
;
State Key Laboratory of Precision and Intelligent Chemistry(精准与智能化学国家重点实验室)
机构
*
School of Computer Science, Peking University(北京大学计算机科学学院)
;
School of Computer Science, Zhejiang University(浙江大学计算机科学学院)
;
School of Cyber Science and Engineering, Wuhan University(武汉大学网络科学与工程学院)
;
School of Information Science and Technology, Tibet University(西藏大学信息科学与技术学院)