arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-10 至 2025-10-10 共收录 16 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 16 篇

2509.14830 2025-10-10 cs.CV cs.AI cs.LG 85%

ProtoMedX: Towards Explainable Multi-Modal Prototype Learning for Bone Health Classification

Alvaro Lopez Pellicer, Andre Mariucci, Plamen Angelov, Marwan Bukhari, Jemma G. Kerns

机构 * School of Computing and Communications(计算与通讯学院) Lancaster Medical School(兰卡斯特医学学院)

专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract,comments);分类 cs.CV、cs.AI

Comments ICCV 2025 (PHAROS-AFE-AIMI: Adaptation, Fairness, and Explainability in Medical Imaging). 8 pages, 5 figures, 4 tables. Keywords: multi-modal, multimodal, prototype learning, explainable AI, interpretable models, case-based reasoning, medical imaging, DEXA, bone health, osteoporosis, osteopenia, diagnosis, classification, clustering

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06871 2025-10-10 cs.LG cs.CV 83%

SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models

Huahui Yi, Kun Wang, Qiankun Li, Miao Yu, Liang Lin, Gongli Xi, Hao Wu, Xuming Hu, Kang Li, Yang Liu

机构 * West China Biomedical Big Data Center, West China Hospital, SCU(西昌生物医学大数据中心、西昌医院、SCU) NTU USTC TeleAI, China Telecom(TeleAI、中国电信) BUPT Tsinghua University(清华大学) HKUST(Guangzhou)(HKUST(广州))

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06019 2025-10-10 cs.CV cs.AI eess.IV eess.SP 81%

BRIGHT: A globally distributed multimodal building damage assessment dataset with very-high-resolution for all-weather disaster response

Hongruixuan Chen, Jian Song, Olivier Dietrich, Clifford Broni-Bediako, Weihao Xuan, Junjue Wang, Xinlei Shao, Yimin Wei, Junshi Xia, Cuiling Lan, Konrad Schindler, Naoto Yokoya

机构 * Graduate School of Frontier Sciences, The University of Tokyo(东京大学前沿科学研究生院) RIKEN Center for Advanced Intelligence Project (AIP), RIKEN(日本理化学研究院先进智能项目中心) Department of Photogrammetry and Remote Sensing, ETH Zürich(苏黎世联邦理工学院测绘与遥感系) Microsoft Research Asia(微软亚洲研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10894 2025-10-10 cs.CV 79%

MAESTRO: Masked AutoEncoders for Multimodal, Multitemporal, and Multispectral Earth Observation Data

Antoine Labatie, Michael Vaccaro, Nina Lardiere, Anatol Garioud, Nicolas Gonthier

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13468 2025-10-10 cs.RO cs.AI cs.HC 79%

ERR@HRI 2.0 Challenge: Multimodal Detection of Errors and Failures in Human-Robot Conversations

Shiye Cao, Maia Stiber, Amama Mahmood, Maria Teresa Parreira, Wendy Ju, Micol Spitale, Hatice Gunes, Chien-Ming Huang

机构 * Johns Hopkins University(约翰霍普金斯大学) Microsoft Research(微软研究院) Cornell University(康奈尔大学) University of Cambridge(剑桥大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07636 2025-10-10 cs.CV 79%

PIT-QMM: A Large Multimodal Model For No-Reference Point Cloud Quality Assessment

Shashank Gupta, Gregoire Phillips, Alan C. Bovik

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Ericsson Research(爱立信研究)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Oral presentation at ICIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01444 2025-10-10 cs.CR cs.AI 79%

PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization

Aofan Liu, Lulu Tang, Ting Pan, Yuguo Yin, Bin Wang, Ao Yang

机构 * School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments Accepted to IEEE International Conference on Multimedia and Expo (ICME) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07325 2025-10-10 cs.LG cs.NE 78%

A Modality-Aware Cooperative Co-Evolutionary Framework for Multimodal Graph Neural Architecture Search

Sixuan Wang, Jiao Yin, Jinli Cao, Mingjian Tang, Yong-Feng Ge

机构 * Department of Computer Science and Information Technology, La Trobe University(计算机科学与信息技术系,拉特罗布大学) Institute for Sustainable Industries and Liveable Cities, Victoria University(可持续产业与宜居城市研究所,维多利亚大学)

专题命中 多模态评测 :multimodal(title,abstract)

Comments 11 pages, 6 figures. This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08011 2025-10-10 cs.CV cs.CL 73%

Play to Generalize: Learning to Reason Through Game Play

Yunfei Xie, Yinsong Ma, Shiyi Lan, Alan Yuille, Junfei Xiao, Chen Wei

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.CL

Comments Project Page: https://yunfeixie233.github.io/ViGaL/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16746 2025-10-10 cs.CV cs.CL cs.LG 66%

Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning

Ang Li, Charles Wang, Deqing Fu, Kaiyu Yue, Zikui Cai, Wang Bill Zhu, Ollie Liu, Peng Guo, Willie Neiswanger, Furong Huang, Tom Goldstein, Micah Goldblum

机构 * Columbia University(哥伦比亚大学) University of Maryland(马里兰大学) University of Southern California(南加州大学) New York University(纽约大学)

专题命中 多模态评测 :multimodal(abstract,comments);分类 cs.CV、cs.CL

Comments dataset link: https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08202 2025-10-10 cs.HC cs.AI cs.CL cs.ET 62%

Sentiment Matters: An Analysis of 200 Human-SAV Interactions

Lirui Guo, Michael G. Burke, Wynita M. Griggs

机构 * Department of Civil and Environmental Engineering, Monash University(莫纳什大学土木与环境工程系) Department of Electrical and Computer Systems Engineering, Monash University(莫纳什大学电子与计算机系统工程系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted for presentation at IEEE ITSC 2025 and for publication in its Proceedings. \c{opyright} 2025 IEEE. Personal use permitted; other uses require permission from IEEE, including reprinting, republishing, or reuse of any copyrighted component of this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13546 2025-10-10 cs.CV 57%

GazeProphet: Software-Only Gaze Prediction for VR Foveated Rendering

Farhaan Ebadulla, Chiraag Mudlapur, Gaurav BV

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08022 2025-10-10 cs.RO cs.AI 57%

FastUMI-100K: Advancing Data-driven Robotic Manipulation with a Large-scale UMI-style Dataset

Kehui Liu, Zhongjie Jia, Yang Li, Zhaxizhuoma, Pengan Chen, Song Liu, Xin Liu, Pingrui Zhang, Haoming Song, Xinyi Ye, Nieqing Cao, Zhigang Wang, Jia Zeng, Dong Wang, Yan Ding, Bin Zhao, Xuelong Li

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Northwestern Polytechnical University(西北工业大学) Shanghai Jiao Tong University(上海交通大学) TongJi University(同济大学) Xi’an Jiaotong-Liverpool University(西安交通大学-利物浦大学) Suzhou OneStar Robotics Corp Ltd(苏州OneStar机器人有限公司) Institute of Artificial Intelligence, China Telecom Corp Ltd(中国电信人工智能研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04097 2025-10-10 cs.AI 57%

WebRenderBench: Enhancing Web Interface Generation through Layout-Style Consistency and Reinforcement Learning

Peichao Lai, Jinhui Zhuang, Kexuan Zhang, Ningchang Xiong, Shengjie Wang, Yanwei Xu, Chong Chen, Yilei Wang, Bin Cui

机构 * Peking University(北京大学) Xiamen Huaxia University(厦门华夏大学) Fuzhou University(福州市大学) City University of Hong Kong(香港城市大学) Huawei Cloud BU(华为云业务单元)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01571 2025-10-10 cs.CV 57%

PainFormer: a Vision Foundation Model for Automatic Pain Assessment

Stefanos Gkikas, Raul Fernandez Rojas, Manolis Tsiknakis

机构 * Hellenic Mediterranean University, Department of Electrical and Computer Engineering(希腊地中海大学电子与计算机工程系) Institute of Computer Science, Foundation for Research & Technology-Hellas(希腊研究所计算机科学研究所) University of Canberra, Faculty of Science and Technology(堪培拉大学科学与技术学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Journal ref IEEE Transactions on Affective Computing; 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08106 2025-10-10 cs.RO 50%

Beyond hospital reach: Autonomous lightweight ultrasound robot for liver sonography

Zihan Li, Yixiao Xu, Lei Zhang, Taiyu Han, Xinshan Yang, Yingni Wang, Mingxuan Liu, Shenghai Xin, Linxun Liu, Hongen Liao, Guochen Ning

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏