arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-12 至 2025-08-12 共收录 24 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 24 篇

2508.07766 2025-08-12 cs.CV cs.AI 84%

UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models

Jinke Li, Jiarui Yu, Chenxing Wei, Hande Dong, Qiang Lin, Liangjing Yang, Zhicai Wang, Yanbin Hao

机构 * Zhejiang University(浙江大学) Tencent(腾讯) Shenzhen University(深圳大学) Hefei University of Technology(合肥工业大学)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted at ACM MM 2025 Dataset Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08093 2025-08-12 cs.CV cs.LG cs.MM eess.AS 82%

MDD-Net: Multimodal Depression Detection through Mutual Transformer

Md Rezwanul Haque, Md. Milon Islam, S M Taslim Uddin Raju, Hamdi Altaheri, Lobna Nassar, Fakhri Karray

机构 * Centre for Pattern Analysis and Machine Intelligence, Department of Electrical and Computer Engineering, University of Waterloo(模式分析与机器智能中心,电气与计算机工程系,滑铁卢大学) School of Engineering and Computing, Department of Computer Science and Engineering, American University of Ras Al Khaimah(工程与计算学院,计算机科学与工程系,阿联酋拉线哈姆市美国大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted for the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17632 2025-08-12 cs.AI cs.CV cs.MM 82%

D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal Guidance

Renyang Liu, Ziyu Lyu, Wei Zhou, See-Kiong Ng

机构 * School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学信息科学与技术学院(深圳校区)) Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究所) College of Modern Engineering and the Engineering Research Center of Cyberspace, Yunnan University(云南大学现代工程学院及空天信息工程研究中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06272 2025-08-12 cs.CV cs.AI 81%

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance

Zhang Li, Biao Yang, Qiang Liu, Shuo Zhang, Zhiyin Ma, Liang Yin, Linger Deng, Yabo Sun, Yuliang Liu, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07628 2025-08-12 cs.AI 79%

Multimodal AI Systems for Enhanced Laying Hen Welfare Assessment and Productivity Optimization

Daniel Essien, Suresh Neethirajan

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 66 pages, 7 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07596 2025-08-12 cs.CV 79%

From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users

Shahroz Tariq, Simon S. Woo, Priyanka Singh, Irena Irmalasari, Saakshi Gupta, Dev Gupta

机构 * Sungkyunkwan University, S. Korea(顺天大学) University of Queensland, Australia(昆士兰大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 3 tables, 5 figures, accepted for publicaiton in the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06851 2025-08-12 cs.AI cs.CY 79%

MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams

Pengfei Zhou, Xiaopeng Peng, Fanrui Zhang, Zhaopan Xu, Jiaxin Ai, Yansheng Qiu, Chuanhao Li, Zhen Li, Ming Li, Yukang Feng, Jianwen Sun, Haoquan Zhang, Zizhen Li, Xiaofeng Mao, Zekai Li, Wangbo Zhao, Kai Wang, Xiaojun Chang, Wenqi Shao, Yang You, Kaipeng Zhang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) USTC(中国科学技术大学) RIT(罗切斯特理工学院) HIT(哈尔滨工业大学) WHU(武汉大学) MBZUAI(马克斯·普朗克人工智能研究所) NUS(新加坡国立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 35 pages, 33 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06558 2025-08-12 cs.CV cs.LG 79%

On the effectiveness of multimodal privileged knowledge distillation in two vision transformer based diagnostic applications

Simon Baur, Alexandra Benova, Emilio Dolgener Cantú, Jackie Ma

机构 * Fraunhofer Heinrich-Hertz-Institut(弗劳恩霍夫 Heinrich-Hertz 研究所) Universität Osnabrück(奥斯纳布吕克大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16104 2025-08-12 cs.CY 78%

Beauty and the Bias: Exploring the Impact of Attractiveness on Multimodal Large Language Models

Aditya Gulati, Moreno D'Incà, Nicu Sebe, Bruno Lepri, Nuria Oliver

专题命中 多模态评测 :multimodal(title,abstract)

Comments 39 pages, 4 figures, 33 tables; Accepted for publication at the Eighth AAAI/ACM Conference on AI, Ethics and Society (AIES 2025) (https://www.aies-conference.com/2025/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07818 2025-08-12 cs.CV 70%

Segmenting and Understanding: Region-aware Semantic Attention for Fine-grained Image Quality Assessment with Large Language Models

Chenyue Song, Chen Hui, Haiqi Zhu, Feng Jiang, Yachun Mi, Wei Zhang, Shaohui Liu

专题命中 多模态评测 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07554 2025-08-12 cs.MM 70%

FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding

Xusheng He, Wei Liu, Shanshan Ma, Qian Liu, Chenghao Ma, Jianlong Wu

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14533 2025-08-12 cs.CV 70%

ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding

Shuo Cao, Nan Ma, Jiayang Li, Xiaohui Li, Lihao Shao, Kaiwen Zhu, Yu Zhou, Yuandong Pu, Jiarui Wu, Jiaquan Wang, Bo Qu, Wenhai Wang, Yu Qiao, Dajuin Yao, Yihao Liu

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments 43 pages, 31 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06565 2025-08-12 cs.CV cs.LG 70%

Bridging Brain Connectomes and Clinical Reports for Early Alzheimer's Disease Diagnosis

Jing Zhang, Xiaowei Yu, Minheng Chen, Lu Zhang, Tong Chen, Yan Zhuang, Chao Cao, Yanjun Lyu, Li Su, Tianming Liu, Dajiang Zhu

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07022 2025-08-12 cs.AI cs.CL cs.LG cs.MM 67%

MultiMedEdit: A Scenario-Aware Benchmark for Evaluating Knowledge Editing in Medical VQA

Shengtao Wen, Haodong Chen, Yadong Wang, Zhongying Pan, Xiang Chen, Yu Tian, Bo Qian, Dong Liang, Sheng-Jun Huang

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI、cs.MM

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07714 2025-08-12 cs.CV cs.AI cs.ET 62%

DoorDet: Semi-Automated Multi-Class Door Detection Dataset via Object Detection and Large Language Models

Licheng Zhang, Bach Le, Naveed Akhtar, Tuan Ngo

机构 * The University of Melbourne(墨尔本大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20988 2025-08-12 cs.AI cs.CL 62%

Reviewing Clinical Knowledge in Medical Large Language Models: Training and Beyond

Qiyuan Li, Haijiang Liu, Caicai Guo, Chao Gao, Deyu Chen, Meng Wang, Feng Gao, Frank van Harmelen, Jinguang Gu

机构 * School of Computer Science and Technology, Wuhan University of Science and Technology(计算机科学与技术学院,武汉科技大学) Hubei Province Key Laboratory of Intelligent Information Processing and Real-time Industrial System(湖北省智能信息处理与实时工业系统重点实验室) School of Computer Science and Technology, Huazhong University of Science and Technology(计算机科学与技术学院,华中科技大学) School of Cyber Science and Engineering, Wuhan University(网络科学与工程学院,武汉大学) Department of Computer Science, Vrije Universiteit Amsterdam(计算机科学系,阿姆斯特丹自由大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted for publication in Knowledge-Based Systems. The arXiv version is the pre-peer-review preprint, and the final published version is not available here due to publisher policy

Journal ref Knowledge-Based Systems, 114215(2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08252 2025-08-12 cs.CV 57%

ReferSplat: Referring Segmentation in 3D Gaussian Splatting

Shuting He, Guangquan Jie, Changshuo Wang, Yun Zhou, Shuming Hu, Guanbin Li, Henghui Ding

机构 * MoE Key Laboratory of Interdisciplinary Research of Computation and Economics(跨学科计算与经济学研究实验室) Shanghai University of Finance(上海财经大学) Institute of Big Data, College of Computer Science and Artificial Intelligence(大数据研究所,计算机科学与人工智能学院) Fudan University, Shanghai, China(复旦大学,上海,中国) Nanyang Technological University, Singapore(南洋理工大学,新加坡) Sun Yat-sen University, Guangzhou, China(中山大学,广州,中国)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments ICML 2025 Oral, Code: https://github.com/heshuting555/ReferSplat

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07838 2025-08-12 cs.CV 57%

CBDES MoE: Hierarchically Decoupled Mixture-of-Experts for Functional Modules in Autonomous Driving

Qi Xiang, Kunsong Shi, Zhigui Lin, Lei He

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07812 2025-08-12 cs.CV 57%

Semi-supervised Multiscale Matching for SAR-Optical Image

Jingze Gai, Changchun Li

机构 * Nanyang Technological University(南洋理工大学) Jilin University(吉林大学)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV

Comments 15 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05016 2025-08-12 cs.CV eess.IV 57%

AU-IQA: A Benchmark Dataset for Perceptual Quality Assessment of AI-Enhanced User-Generated Content

Shushi Wang, Chunyi Li, Zicheng Zhang, Han Zhou, Wei Dong, Jun Chen, Guangtao Zhai, Xiaohong Liu

机构 * Shanghai Jiao Tong University(上海交通大学) McMaster University(麦斯特大学) Suzhou Key Laboratory of Artificial Intelligence(苏州人工智能重点实验室)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted by ACMMM 2025 Datasets Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06624 2025-08-12 cs.CV 57%

VL-MedGuide: A Visual-Linguistic Large Model for Intelligent and Explainable Skin Disease Auxiliary Diagnosis

Kexin Yu, Zihan Xu, Jialei Xie, Carter Adams

机构 * Jiangsu Ocean University(江苏海洋大学) Federal University of Bahia(巴伊亚联邦大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01416 2025-08-12 cs.RO cs.CV 57%

UniCalib: Targetless LiDAR-Camera Calibration via Probabilistic Flow on Unified Depth Representations

Shu Han, Xubo Zhu, Ji Wu, Ximeng Cai, Wen Yang, Huai Yu, Gui-Song Xia

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) School of Electronic Information, Wuhan University(武汉大学电子信息学院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV

Comments 8 pages,5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04305 2025-08-12 eess.IV 50%

Edge2Prompt: Modality-Agnostic Model for Out-of-Distribution Liver Segmentation

Nathan Hollet, Oumeymah Cherkaoui, Philippe C. Cattin, Sidaty El Hadramy

专题命中 多模态评测 :multi-modal(abstract)

Comments 8 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06841 2025-08-12 cs.NE 50%

Memory Enhanced Fractional-Order Dung Beetle Optimization for Photovoltaic Parameter Identification

Yiwei Li, Zhihua Allen-Zhao, Yuncheng Xu, Sanyang Liu

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏