arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-12 至 2025-11-12 共收录 11 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 11 篇

2511.08263 2025-11-12 cs.CV cs.AI 84%

ImagebindDC: Compressing Multi-modal Data with Imagebind-based Condensation

Yue Min, Shaobo Wang, Jiaze Li, Tianle Niu, Junxin Fan, Yongliang Miao, Lijin Yang, Linfeng Zhang

机构 * EPIC Lab, SJTU(上海交通大学EPIC实验室) Bosch Corporate Research Asia Pacific(博世亚太公司研究部) HKUST(香港科技大学)

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

Comments AAAI 2026, 18 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07812 2025-11-12 cs.CV 83%

Revisiting MLLM Based Image Quality Assessment: Errors and Remedy

Zhenchen Tang, Songlin Yang, Bo Peng, Zichuan Wang, Jing Dong

专题命中 多模态评测 :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13231 2025-11-12 cs.CV cs.AI 81%

WildFireCan-MMD: A Multimodal Dataset for Classification of User-Generated Content During Wildfires in Canada

Braeden Sherritt, Isar Nejadgholi, Efstratios Aivaliotis, Khaled Mslmani, Marzieh Amini

机构 * Carleton University(卡尔顿大学) National Research Council Canada(加拿大国家研究委员会)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22262 2025-11-12 cs.CV 79%

UniMapGen: A Generative Framework for Large-Scale Map Construction from Multi-modal Data

Yujian Yuan, Changjie Wu, Xinyuan Chang, Sijin Wang, Hang Zhang, Shiyi Liang, Shuang Zeng, Mu Xu, Ning Guo

机构 * Alibaba Group(阿里巴巴集团)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments AAAI2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13861 2025-11-12 cs.HC cs.CL cs.MA 79%

3MDBench: Medical Multimodal Multi-agent Dialogue Benchmark

Ivan Sviridov, Amina Miftakhova, Artemiy Tereshchenko, Galina Zubkova, Pavel Blinov, Andrey Savchenko

机构 * Sber AI Lab(Sber AI实验室) HSE University(俄罗斯高等经济大学) ISP RAS Research Center for Trusted Artificial Intelligence(俄罗斯科学院信息与系统问题研究所可信人工智能研究中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

Comments EMNLP 25 (main)

Journal ref https://aclanthology.org/2025.emnlp-main.1353/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07777 2025-11-12 eess.SP 78%

A Causal-Guided Multimodal Large Language Model for Generalized Power System Time-Series Data Analytics

Zhenghao Zhou, Yiyan Li, Xinjie Yu, Runlong Liu, Zelin Guo, Zheng Yan, Mo-Yuen Chow, Yuqi Yang, Yang Xu

专题命中 多模态评测 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08133 2025-11-12 cs.CV cs.AI 73%

OTSNet: A Neurocognitive-Inspired Observation-Thinking-Spelling Pipeline for Scene Text Recognition

Lixu Sun, Nurmemet Yolwas, Wushour Silamu

机构 * School of Computer Science and Technology(计算机科学与技术学院) School of Computer Science(计算机科学学院) Xinjiang University(新疆大学)

专题命中 多模态评测 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14686 2025-11-12 cs.CV 70%

From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition

Chen Cai, Tianyi Liu, Jianjun Gao, Wenyang Liu, Kejun Wu, Ruoyu Wang, Yi Wang, Soo Chin Liew

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) Huazhong University of Science and Technology(华中科技大学) The Hong Kong Polytechnic University(香港理工大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22340 2025-11-12 cs.AI cs.CL cs.CV cs.LG 67%

DynaSolidGeo: A Dynamic Benchmark for Genuine Spatial Mathematical Reasoning of VLMs in Solid Geometry

Changti Wu, Shijie Lian, Zihao Liu, Lei Zhang, Laurence Tianruo Yang, Kai Chen

机构 * East China Normal University(华东师范大学) Zhongguancun Academy(中关村学院) Huazhong University of Science and Technology(华中科技大学) Peking University(北京大学) Zhengzhou University(郑州大学) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments The code and dataset are available at \href{https://zgca-ai4edu.github.io/DynaSolidGeo/}{DynaSolidGeo}

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08140 2025-11-12 cs.CV 57%

PEOD: A Pixel-Aligned Event-RGB Benchmark for Object Detection under Challenging Conditions

Luoping Cui, Hanqing Liu, Mingjie Liu, Endian Lin, Donghong Jiang, Yuhao Wang, Chuang Zhu

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07897 2025-11-12 cs.AI cs.LG 57%

Data Descriptions from Large Language Models with Influence Estimation

Chaeri Kim, Jaeyeon Bae, Taehwan Kim

机构 * Ulsan National Institute of Science and Technology(UNIST)(乌山国立科学技术研究院)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.AI

Journal ref Published in EMNLP 2025, check our project on this https URL : https://github.com/kimchaeri/Data-Descriptions-from-Large-Language-Models-with-Influence-Estimation

详情

展开后加载摘要…

URL PDF HTML 收藏