MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation
Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu, Xiaoxue Gao, Bryan Chen Zhengyu Tan, Bowei Zou, Chang Liu, Yujia Hu, Xing Xie, Xiaoyuan Yi, Jing Yao, Chaojun Wang, Long Li, Rui Liu, Huiyao Liu, Koji Inoue, Ryuichi Sumida, Tatsuya Kawahara, Fan Xu, Lingyu Ye, Wei Tian, Dongjun Kim, Jimin Jung, Jaehyung Seo, Nadya Yuki Wangsajaya, Pham Minh Duc, Ojasva Saxena, Palash Nandi, Xiyan Tao, Wiwik Karlina, Tuan Luong, Keertana Arun Vasan, Roy Ka-Wei Lee, Nancy F. Chen
机构
*
Singapore University of Technology and Design(新加坡科技设计大学)
;
Agency for Science, Technology and Research, Singapore(新加坡科技研究局)
;
Indian Institute of Technology Delhi(印度理工学院德里分校)
;
Alibaba DAMO Academy(阿里巴巴达摩院)
;
Microsoft Research Asia(微软亚洲研究院)
;
Shanghai University of Finance and Economics(上海财经大学)
;
Inner Mongolia University(内蒙古大学)
;
Kyoto University(京都大学)
;
Jiangxi Normal University(江西师范大学)
;
Korea University(韩国大学)
;
Nanyang Technological University(南洋理工大学)
机构
*
Chongqing University(重庆大学)
;
Independent Researcher(独立研究者)
;
University of the Chinese Academy of Sciences(中国科学院大学)
;
Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航天信息研究所)
专题命中
视觉问答
:multimodal large language model(abstract)
CommentsThe paper is currently under investigation regarding concerns of potential academic misconduct. While the investigation is ongoing, the authors have voluntarily requested to withdraw the manuscript
Unleashing Perception-Time Scaling to Multimodal Reasoning Models
Yifan Li, Zhenghao Chen, Ziheng Wu, Kun Zhou, Ruipu Luo, Can Zhang, Zhentao He, Yufei Zhan, Wayne Xin Zhao, Minghui Qiu
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学北京校区人工智能学院)
;
Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大模型与智能治理重点实验室)
;
ByteDance(字节跳动)
;
University of California, San Diego(加州大学圣地亚哥分校)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
RadVLM: A Multitask Conversational Vision-Language Model for Radiology
Nicolas Deperrois, Hidetoshi Matsuo, Samuel Ruipérez-Campillo, Moritz Vandenhirtz, Sonia Laguna, Alain Ryser, Koji Fujimoto, Mizuho Nishio, Thomas M. Sutter, Julia E. Vogt, Jonas Kluckert, Thomas Frauenfelder, Christian Blüthgen, Farhad Nooralahzadeh, Michael Krauthammer
机构
*
Department of Radiology, Kobe University(金泽大学放射科)
;
Department of Computer Science, ETH Zurich(苏黎世联邦理工学院计算机科学系)
;
Department of Advanced Imaging in Medical Magnetic Resonance, Kyoto University(京都大学医学磁共振高级成像部门)
;
Department of Quantitative Biomedicine, University of Zurich(苏黎世大学定量生物医学系)
;
Diagnostic and Interventional Radiology, University Hospital Zurich(苏黎世大学医院诊断与介入放射科)
CommentsWe believe that the contribution of this paper is not enough, so we integrated it into another new paper. The arXiv ID of the new paper is arXiv:2510.01795
机构
*
Zhejiang Gongshang University(浙江工商大学)
;
Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative(宁波空间智能与数字衍生关键实验室)
;
Institute of Digital Twin, Eastern Institute of Technology, Ningbo(数字孪生研究院,东部技术研究所,宁波)
;
Meituan Inc.(美团公司)
;
National University of Singapore(新加坡国立大学)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Re-Identifying Kākā with AI-Automated Video Key Frame Extraction
Paula Maddigan, Andrew Lensen, Rachael C. Shaw
机构
*
Centre for Data Science and Artificial Intelligence, and School of Engineering and Computer Science(数据科学与人工智能中心,工程与计算机科学学院)
;
Victoria University of Wellington(惠灵顿维多利亚大学)
;
School of Biological Sciences(生物科学学院)
机构
*
Bytedance(字节跳动)
;
Peking University(北京大学)
;
Peking University People’s Hospital(北京大学人民医院)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
The Chinese University of Hong Kong(香港中文大学)
专题命中
GUI与屏幕智能体
:multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving
Haibo Hu, Lianming Huang, Xinyu Wang, Yufei Cui, Shangyu Wu, Nan Guan, Chun Jason Xue
机构
*
Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系)
;
Department of Computer Science, McGill University(麦吉尔大学计算机科学系)
;
Department of Computer Science, Mohamed bin Zayed University of Artificial Intelligence(马尔代夫穆罕默德· bin·扎耶德人工智能大学计算机科学系)
Mem4Nav: Boosting Vision-and-Language Navigation in Urban Environments with a Hierarchical Spatial-Cognition Long-Short Memory System
Lixuan He, Haoyu Dong, Zhenxing Chen, Yangcheng Yu, Jie Feng, Yong Li
专题命中
GUI与屏幕智能体
:MLLM(abstract);分类 cs.CV、cs.AI
CommentsThe paper is currently under investigation regarding concerns of potential academic misconduct. While the investigation is ongoing, the authors have voluntarily requested to withdraw the manuscript