arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-10-07 至 2025-10-07 共收录 22 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 22 篇

2506.01713 2025-10-07 cs.CL 83%

SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning

Zhongwei Wan, Zhihao Dou, Che Liu, Yu Zhang, Dongfei Cui, Qinjian Zhao, Hui Shen, Jing Xiong, Yi Xin, Yifan Jiang, Chaofan Tao, Yangfan He, Mi Zhang, Shen Yan

机构 * The Ohio State University(俄亥俄州立大学)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00883 2025-10-07 cs.CL 83%

Improve MLLM Benchmark Efficiency through Interview

Farong Wen, Yijin Guo, Junying Wang, Jiaohao Xiao, Yingjie Zhou, Ye Shen, Qi Jia, Chunyi Li, Zicheng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 多模态评测 :MLLM(title,abstract);multimodal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.11625 2025-10-07 cs.CL cs.AI cs.CV cs.LG 82%

MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering

Varun Srivastava, Fan Lei, Srija Mukhopadhyay, Vivek Gupta, Ross Maciejewski

机构 * School of Computing and Augmented Intelligence(计算与增强智能学院) Arizona State University(亚利桑那州立大学) Department of Computer Science(计算机科学系) International Institute of Information Technology(国际信息科技研究所)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

Comments Published as a conference paper at COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03878 2025-10-07 cs.CV cs.AI 81%

Multi-Modal Oral Cancer Detection Using Weighted Ensemble Convolutional Neural Networks

Ajo Babu George, Sreehari J R Ajo Babu George, Sreehari J R Ajo Babu George, Sreehari J R

机构 * Dicemed

专题命中 多模态评测 :multi-modal(title);multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07640 2025-10-07 cs.MM cs.AI 81%

Synthesizing Sentiment-Controlled Feedback For Multimodal Text and Image Data

Puneet Kumar, Sarthak Malik, Balasubramanian Raman, Xiaobai Li

机构 * Center for Machine Vision and Signal Analysis, University of Oulu(机器视觉与信号分析中心,奥卢大学) Indian Institute of Technology Roorkee(印度理工学院罗尔基分校) State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07048 2025-10-07 cs.CV 80%

Comprehensive Evaluation of Large Multimodal Models for Nutrition Analysis: A New Benchmark Enriched with Contextual Metadata

Bruce Coburn, Jiangpeng He, Megan E. Rollo, Satvinder S. Dhaliwal, Deborah A. Kerr, Fengqing Zhu

机构 * Purdue University(普渡大学) Indiana University(印第安纳大学) Curtin University(Curtin大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments The extended full version of the accepted paper in 2025 IEEE BHI conference with title: Evaluating Large Multimodal Models for Nutrition Analysis: A New Benchmark Enriched with Contextual Metadata. Dataset is available at: https://skynet.ecn.purdue.edu/~coburn6/ACETADA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04257 2025-10-07 cs.CR cs.AI 79%

AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents

Yanjie Li, Yiming Cao, Dong Wang, Bin Xiao

机构 * Computing Department of Hong Kong Polytechnic University(香港理工大学计算机系) Computing Department, The Hong Kong Polytechnic University(香港理工大学计算机系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 13 pages, 8 figures. Submitted to IEEE Transactions on Information Forensics & Security

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03965 2025-10-07 cs.MM 79%

FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction

Dong Shu, Yanguang Liu, Huopu Zhang, Mengnan Du

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12230 2025-10-07 cs.CV cs.AI cs.RO 76%

LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps

Yihao Wang, Raphael Memmesheimer, Sven Behnke

机构 * Autonomous Intelligent Systems, Computer Science Institute VI, Center for Robotics, University of Bonn, Germany(自主智能系统,计算机科学研究所第六研究所,机器人中心,波恩大学,德国)

专题命中 多模态评测 :multimodal(title);分类 cs.CV、cs.AI

Comments 12 pages, 4 figures, 2 tables, 19th International Conference on Intelligent Autonomous Systems (IAS), Genoa, Italy, June 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04438 2025-10-07 cs.CV cs.CL 62%

The Telephone Game: Evaluating Semantic Drift in Unified Models

Sabbir Mollah, Rohit Gupta, Sirnam Swetha, Qingyang Liu, Ahnaf Munir, Mubarak Shah

机构 * Center For Research in Computer Vision, University of Central Florida, USA(计算机视觉研究中心,中央佛罗里达大学)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03555 2025-10-07 cs.CV cs.AI 62%

GAS-MIL: Group-Aggregative Selection Multi-Instance Learning for Ensemble of Foundation Models in Digital Pathology Image Analysis

Peiran Quan, Zifan Gu, Zhuo Zhao, Qin Zhou, Donghan M. Yang, Ruichen Rong, Yang Xie, Guanghua Xiao

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23631 2025-10-07 cs.HC cs.AI cs.CL 62%

Human Empathy as Encoder: AI-Assisted Depression Assessment in Special Education

Boning Zhao, Xinnuo Li, Yutong Hu

机构 * Tandon School of Engineering New York University New York, USA College of Arts \& Science New York University New York, USA

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 6 figures, ACII 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19099 2025-10-07 cs.AI physics.ed-ph physics.pop-ph 57%

SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning

Kun Xiang, Heng Li, Terry Jingchen Zhang, Yinya Huang, Zirong Liu, Peixin Qu, Jixi He, Jiaqi Chen, Yu-Jie Yuan, Jianhua Han, Hang Xu, Hanhui Li, Mrinmaya Sachan, Xiaodan Liang

机构 * Sun Yat-sen University(中山大学) ETH Zurich(苏黎世联邦理工学院) Huawei Noah’s Ark Lab(华为诺亚实验室) The University of Hong Kong(香港大学) ETH AI Center(苏黎世人工智能中心)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments 46 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11334 2025-10-07 cs.AI 57%

Program Synthesis Benchmark for Visual Programming in XLogoOnline Environment

Chao Wen, Jacqueline Staub, Adish Singla

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

Comments ACL'25 paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03874 2025-10-07 cs.CV 57%

DHQA-4D: Perceptual Quality Assessment of Dynamic 4D Digital Human

Yunhao Li, Sijing Wu, Yucheng Zhu, Huiyu Duan, Zicheng Zhang, Guangtao Zhai

机构 * Institute of Image Communication and Network Engineering(图像通信与网络工程研究所) Shanghai Jiao Tong University(上海交通大学) USC-SJTU Institute of Cultural and Creative Industry(USC-SJTU文化产业与创意产业研究院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03863 2025-10-07 cs.AI cs.CR 57%

Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine Differentiation

Arina Kharlamova, Bowei He, Chen Ma, Xue Liu

专题命中 多模态评测 :multi-modal(abstract);分类 cs.AI

Comments Submitted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01844 2025-10-07 cs.AI 57%

Towards Generalizable Context-aware Anomaly Detection: A Large-scale Benchmark in Cloud Environments

Xinkai Zou, Xuan Jiang, Ruikai Huang, Haoze He, Parv Kapoor, Hongrui Wu, Yibo Wang, Jian Sha, Xiongbo Shi, Zixun Huang, Jinhua Zhao

机构 * UC San Diego(加州大学圣地亚哥分校) Massachusetts Institute of Technology(麻省理工学院) Georgia Institute of Technology(佐治亚理工学院) Carnegie Mellon University(卡内基梅隆大学) Tongji University(同济大学) Tsinghua University(清华大学) University of Pennsylvania(宾夕法尼亚大学)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09255 2025-10-07 cs.CV 57%

FVQ: A Large-Scale Dataset and an LMM-based Method for Face Video Quality Assessment

Sijing Wu, Yunhao Li, Ziwen Xu, Yixuan Gao, Huiyu Duan, Wei Sun, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) East China Normal University(华东师范大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025. Project page: https://github.com/wsj-sjtu/FVQ

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03376 2025-10-07 cs.CV eess.IV 57%

Visual Language Model as a Judge for Object Detection in Industrial Diagrams

Sanjukta Ghosh

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Pre-review version submitted to IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04183 2025-10-07 cs.NI 50%

Dynamic Adaptive Federated Learning for mmWave Sector Selection

Lucas Pacheco, Torsten Braun, Kaushik Chowdhury, Denis Rosário, Batool Salehi, Eduardo Cerqueira

专题命中 多模态评测 :multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04122 2025-10-07 cs.HC 50%

Wrist2Finger: Sensing Fingertip Force for Force-Aware Hand Interaction with a Ring-Watch Wearable

Yingjing Xiao, Zhichao Huang, Junbin Ren, Haichuan Song, Yang Gao, Yuting Bai, Zhanpeng Jin

专题命中 多模态评测 :cross-modal(abstract)

Comments 15 pages, 13 figures. Accepted at UIST 2025 (ACM Symposium on User Interface Software and Technology). Yingjing Xiao and Zhichao Huang contributed equally. Corresponding author: Yang Gao (gaoyang@cs.ecnu.edu.cn)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03815 2025-10-07 eess.SY cs.LG cs.SY eess.SP 50%

A Trustworthy Industrial Fault Diagnosis Architecture Integrating Probabilistic Models and Large Language Models

Yue wu

专题命中 多模态评测 :multimodal(abstract)

Comments 1tables,6 figs,11pages

详情

展开后加载摘要…

URL PDF HTML 收藏