arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-07-29 至 2025-07-29 共收录 21 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 21 篇

2411.17776 2025-07-29 cs.CV cs.MM 84%

Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search

Shuyu Yang, Yaxiong Wang, Li Zhu, Zhedong Zheng

机构 * Xi’an Jiaotong University(西安交通大学) Hefei University of Technology(合肥工业大学) University of Macau(澳门大学)

专题命中 多模态评测 :image-text(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03328 2025-07-29 cs.CV cs.AI cs.NE 84%

Visual Enumeration Remains Challenging for Multimodal Generative AI

Alberto Testolin, Kuinan Hou, Marco Zorzi

机构 * Department of General Psychology and Department of Mathematics University of Padova(帕多瓦大学心理学系和数学系) Department of General Psychology University of Padova(帕多瓦大学心理学系) Department of General Psychology and Padova Neuroscience Center University of Padova(帕多瓦大学心理学系和帕多瓦神经科学中心) IRCSS San Camillo Hospital, Venice-Lido(威尼斯利多医院IRCSS桑卡莫医院)

专题命中 多模态评测 :multimodal(title,abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20388 2025-07-29 cs.CV 83%

ModalFormer: Multimodal Transformer for Low-Light Image Enhancement

Alexandru Brateanu, Raul Balmez, Ciprian Orhei, Codruta Ancuti, Cosmin Ancuti

机构 * Department of Computer Science University of Manchester(计算机科学系曼彻斯特大学) Department of Computer and Information Technology Politehnica University of Timisoara(计算机与信息科技系蒂米什瓦拉工业大学)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19525 2025-07-29 cs.LG cs.AI 83%

MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs

Chenchen Zhao, Zhengyuan Shi, Xiangyu Wen, Chengjie Liu, Yi Liu, Yunhao Zhou, Yuxiang Zhao, Hefei Feng, Yinan Zhu, Gwok-Waa Wan, Xin Cheng, Weiyu Chen, Yongqi Fu, Chujie Chen, Chenhao Xue, Guangyu Sun, Ying Wang, Yibo Lin, Jun Yang, Ning Xu, Xi Wang, Qiang Xu

机构 * Department of Computer Science and Engineering, The Chinese University of Hong Kong(中国香港中文大学计算机科学与工程系) School of Electronic Science and Engineering, Nanjing University(南京大学电子科学与工程学院) School of Integrated Circuits, Peking University(北京大学集成电路学院) School of Intergrated Circuits, Southeast University(东南大学集成电路学院) School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Department of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术系) National Center of Technology Innovation for EDA(EDA技术创新国家中心)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments 10 pages, 1 figure, 5 tables. To appear in ICCAD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20872 2025-07-29 cs.CV cs.AI cs.LG 81%

Not Only Grey Matter: OmniBrain for Robust Multimodal Classification of Alzheimer's Disease

Ahmed Sharshar, Yasser Ashraf, Tameem Bakr, Salma Hassan, Hosam Elgendy, Mohammad Yaqub, Mohsen Guizani

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Published in Third Workshop on Computer Vision for Automated Medical Diagnosis CVAMD 2025 in ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20737 2025-07-29 cs.CV cs.AI cs.HC 81%

Multi-Masked Querying Network for Robust Emotion Recognition from Incomplete Multi-Modal Physiological Signals

Geng-Xin Xu, Xiang Zuo, Ye Li

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518055, China(深圳先进技术研究院,中国科学院,深圳518055,中国) Southern University of Science and Technology, Shenzhen 518055, China(南方科技大学,深圳518055,中国)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments MICCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19969 2025-07-29 cs.CL cs.CV 81%

Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text

Mizanur Rahman, Md Tahmid Rahman Laskar, Shafiq Joty, Enamul Hoque

机构 * York University(约克大学) Dialpad Salesforce AI Research(Salesforce人工智能研究)

专题命中 多模态评测 :multimodal(title);cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19682 2025-07-29 cs.CV cs.AI 81%

DeepJIVE: Learning Joint and Individual Variation Explained from Multimodal Data Using Deep Learning

Matthew Drexler, Benjamin Risk, James J Lah, Suprateek Kundu, Deqiang Qiu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments 26 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20764 2025-07-29 cs.CV 79%

ATR-UMMIM: A Benchmark Dataset for UAV-Based Multimodal Image Registration under Complex Imaging Conditions

Kangcheng Bin, Chen Chen, Ting Hu, Jiahao Qi, Ping Zhong

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20613 2025-07-29 cs.AI cs.LG 79%

Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression

Te Zhang, Yuheng Li, Junxiang Wang, Lujun Li

机构 * University of Michigan(密歇根大学) Johns Hopkins University(约翰霍普金斯大学) Central South university(中南大学) HKGAI

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20503 2025-07-29 cs.LG cs.CL cs.CY 79%

Customize Multi-modal RAI Guardrails with Precedent-based predictions

Cheng-Fu Yang, Thanh Tran, Christos Christodoulopoulos, Weitong Ruan, Rahul Gupta, Kai-Wei Chang

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11780 2025-07-29 cs.CV 79%

Rethinking Multi-Modal Object Detection from the Perspective of Mono-Modality Feature Learning

Tianyi Zhao, Boyang Liu, Yanglei Gao, Yiming Sun, Maoxun Yuan, Xingxing Wei

机构 * Institute of Artificial Intelligence, State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(人工智能研究院、虚拟现实技术与系统国家重点实验室、北京航空航天大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19697 2025-07-29 cs.LG 71%

NAICS-Aware Graph Neural Networks for Large-Scale POI Co-visitation Prediction: A Multi-Modal Dataset and Methodology

Yazeed Alrubyli, Omar Alomeir, Abrar Wafa, Diána Hidvégi, Hend Alrasheed, Mohsen Bahrami

机构 * Prince Sultan University(普森大学) Massachusetts Institute of Technology(麻省理工学院)

专题命中 多模态评测 :multi-modal(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14699 2025-07-29 cs.CV cs.CL cs.LG 62%

Benchmarking Graph Neural Networks for Document Layout Analysis in Public Affairs

Miguel Lopez-Duran, Julian Fierrez, Aythami Morales, Ruben Tolosana, Oscar Delgado-Mohatar, Alvaro Ortigosa

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 15 pages, 2 figures, accepted paper at The Fifth ICDAR International Workshop on Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21033 2025-07-29 cs.CV 57%

GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

Yuhan Wang, Siwei Yang, Bingchen Zhao, Letian Zhang, Qing Liu, Yuyin Zhou, Cihang Xie

机构 * University of California, Santa Cruz(加州大学圣克鲁兹分校) The University of Edinburgh(爱丁堡大学) Adobe Project Page(Adobe项目页面)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10508 2025-07-29 cs.CV 57%

Hoi2Threat: An Interpretable Threat Detection Method for Human Violence Scenarios Guided by Human-Object Interaction

Yuhan Wang, Cheng Liu, Daou Zhang, Zihan Zhao, Jinyang Chen, Purui Dong, Zuyuan Yu, Ziru Wang, Weichao Wu

机构 * organization= School of Mechatronical Engineering, Beijing Institute of Technology , postcode= 100081 , city= Beijing , country= China organization= School of Computer Scienece \& Technology, Beijing Insitute of Technology , postcode= 100081 , city= Beijing , country= China organization= School of Automation, Beijing Insitute of Technology , postcode= 100081 , city= Beijing , country= China organization= School of Information

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20144 2025-07-29 cs.LG cs.AI 57%

Awesome-OL: An Extensible Toolkit for Online Learning

Zeyi Liu, Songqiao Hu, Pengyu Han, Jiaming Liu, Xiao He

机构 * Department of Automation, Tsinghua University(自动化系,清华大学) Department of Computer Science, Beihang University(计算机科学系,北航)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19885 2025-07-29 cs.CL 57%

Zero-shot Performance of Generative AI in Brazilian Portuguese Medical Exam

Cesar Augusto Madid Truyts, Amanda Gomes Rabelo, Gabriel Mesquita de Souza, Daniel Scaldaferri Lages, Adriano Jose Pereira, Uri Adrian Prync Flato, Eduardo Pontes dos Reis, Joaquim Edson Vieira, Paulo Sergio Panse Silveira, Edson Amaro Junior

机构 * Einstein Global Advanced Technologies for Equity(埃因斯坦全球先进科技以公平为宗旨) Hospital Israelita Albert Einstein(埃因斯坦医院) Departamento de Pacientes Graves(重症患者部门) Stanford Center for Artificial Intelligence in Medicine and Imaging(斯坦福大学医学与成像人工智能中心) Departmento de Cirurgia(外科部门) Faculdade de Medicina, Universidade de São Paulo(圣保罗大学医学院) Faculdade Israelita de Ciências da Saúde Albert Einstein(埃因斯坦以色列健康科学学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21745 2025-07-29 cs.CV 57%

3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models

Yuhan Zhang, Mengchen Zhang, Tong Wu, Tengfei Wang, Gordon Wetzstein, Dahua Lin, Ziwei Liu

机构 * Fudan University(复旦大学) Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Stanford University(斯坦福大学) The Chinese University of Hong Kong(香港中文大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室)

专题命中 多模态评测 :MLLM(abstract);分类 cs.CV

Comments Page: https://zyh482.github.io/3DGen-Bench/ ; Code: https://github.com/3DTopia/3DGen-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18796 2025-07-29 cs.HC cs.CY 50%

Artificial Intelligence Can Emulate Human Normative Judgments on Emotional Visual Scenes

Zaira Romeo, Alberto Testolin

专题命中 多模态评测 :multimodal(abstract)

Journal ref Royal Society Open Science, 12(7), 250128 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19493 2025-07-29 cs.HC eess.IV 50%

From Bench to Bedside: A DeepSeek-Powered AI System for Automated Chest Radiograph Interpretation in Clinical Practice

Yaowei Bai, Ruiheng Zhang, Yu Lei, Jingfeng Yao, Shuguang Ju, Chaoyang Wang, Wei Yao, Yiwan Guo, Guilin Zhang, Chao Wan, Qian Yuan, Xuhua Duan, Xinggang Wang, Tao Sun, Yongchao Xu, Chuansheng Zheng, Huangxuan Zhao, Bo Du

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏