arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-05 至 2025-08-05 共收录 18 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 18 篇

2508.02429 2025-08-05 cs.AI cs.LG 85%

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting

Miaosen Luo, Jiesen Long, Zequn Li, Yunying Yang, Yuncheng Jiang, Sijie Mai

机构 * School of Computer Science, South China Normal University(华南师范大学计算机科学学院) School of Information Technology in Education, South China Normal University(华南师范大学教育信息技术学院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01274 2025-08-05 cs.AI cs.CL 84%

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan

Jui-Ming Yao, Bing-Cheng Xie, Sheng-Wei Peng, Hao-Yuan Chen, He-Rong Zheng, Bing-Jia Tan, Peter Shaojui Wang, Shun-Feng Su

机构 * National Taiwan University of Science and Technology(台湾科技大学) University of London(伦敦大学) National Taiwan University(台湾大学)

专题命中 多模态评测 :multimodal(title,abstract);any-to-any(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09071 2025-08-05 cs.CV cs.AI 81%

Segment Any Architectural Facades (SAAF):An automatic segmentation model for building facades, walls and windows based on multimodal semantics guidance

Peilin Li, Jun Yin, Jing Zhong, Ran Luo, Pengyu Zeng, Miao Zhang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01402 2025-08-05 cs.CV 79%

ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models

Chuangchuang Tan, Jinglu Wang, Xiang Ming, Renshuai Tao, Yunchao Wei, Yao Zhao, Yan Lu

机构 * Beijing Jiaotong University(北京交通大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21549 2025-08-05 cs.CV 79%

SiM3D: Single-instance Multiview Multimodal and Multisetup 3D Anomaly Detection Benchmark

Alex Costanzino, Pierluigi Zama Ramirez, Luigi Lella, Matteo Ragaglia, Alessandro Oliva, Giuseppe Lisanti, Luigi Di Stefano

机构 * CVLab, University of Bologna, Italy(博洛尼亚大学计算机视觉实验室) SACMI Imola, Italy(意大利SACMI公司)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2025. Project page: https://alex-costanzino.github.io/SiM3D/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13927 2025-08-05 cs.CV 79%

Multimodal 3D Reasoning Segmentation with Complex Scenes

Xueying Jiang, Lewei Lu, Ling Shao, Shijian Lu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09016 2025-08-05 cs.CV 79%

Cross-Modal Learning for Anomaly Detection in Complex Industrial Process: Methodology and Benchmark

Gaochang Wu, Yapeng Zhang, Lan Deng, Jingxin Zhang, Tianyou Chai

机构 * State Key Laboratory of Synthetical Automation for Process Industries, Northeastern University, Shenyang 110819, P. R. China(合成过程工业系统自动化国家重点实验室,东北大学,沈阳110819,中华人民共和国) School of Software and Electrical Engineering, Swinburne University of Technology, Melbourne, VIC 3122, Australia(软件与电子工程学院,斯威本科技大学,墨尔本,VIC 3122,澳大利亚)

专题命中 多模态评测 :cross-modal(title,abstract);分类 cs.CV

Comments 14 pages, 6 figures, 5 tables. IEEE TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12937 2025-08-05 cs.AI cs.CL cs.CV cs.LG 78%

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Jingyi Zhang, Jiaxing Huang, Huanjin Yao, Shunyu Liu, Xikun Zhang, Shijian Lu, Dacheng Tao

机构 * Nanyang Technological University(南洋理工大学)

专题命中 多模态评测 :multimodal(title);分类 cs.CV、cs.CL、cs.AI

Comments ICCV 2025 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00974 2025-08-05 cs.CV cs.AI cs.LG 62%

ThermoCycleNet: Stereo-based Thermogram Labeling for Model Transition to Cycling

Daniel Andrés López, Vincent Weber, Severin Zentgraf, Barlo Hillen, Perikles Simon, Elmar Schömer

机构 * Institute of Computer Science, Research Group Computational Geometry, Johannes Gutenberg University, Mainz, Germany(计算机科学研究所,计算几何研究组,吉森大学,马因茨,德国) Institute of Sports Science, Department of Sports Medicine, Disease Prevention and Rehabilitation, Johannes Gutenberg University, Mainz, Germany(体育科学研究所,运动医学部门,疾病预防与康复,吉森大学,马因茨,德国) Institute of Occupational, Social and Environmental Medicine, University Medical Center, Mainz, Germany(职业、社会和环境医学研究所,大学医学中心,马因茨,德国)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Presented at IWANN 2025 18th International Work-Conference on Artificial Neural Networks, A Coruña, Spain, 16-18 June, 2025. Book of abstracts: ISBN: 979-13-8752213-1. Funding: Johannes Gutenberg University "Stufe I'': "Start ThermoCycleNet''. Partial funding: Carl-Zeiss-Stiftung: "Multi-dimensionAI'' (CZS-Project number: P2022-08-010)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00893 2025-08-05 cs.CL cs.AI 62%

Affordance Benchmark for MLLMs

Junying Wang, Wenzhe Li, Yalun Wu, Yingji Liang, Yijin Guo, Chunyi Li, Haodong Duan, Zicheng Zhang, Guangtao Zhai

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14240 2025-08-05 cs.CV cs.AI 62%

CityNav: A Large-Scale Dataset for Real-World Aerial Navigation

Jungdae Lee, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto, Daichi Azuma, Yutaka Matsuo, Nakamasa Inoue

机构 * Institute of Science Tokyo(东京科学研究所) The University of Tokyo(东京大学) ATR(ATR公司) National Institute of Informatics(信息研究院)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments ICCV2025. The first two authors are equally contributed. Project page: https://water-cookie.github.io/city-nav-proj/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02645 2025-08-05 cs.CV 61%

Evaluating Variance in Visual Question Answering Benchmarks

Nikitha SR

机构 * Media and Data Science Research Lab, Adobe(媒体与数据科学研究实验室,Adobe)

专题命中 多模态评测 :multimodal(abstract,comments);分类 cs.CV

Comments Accepted in ICCV 2025 Workshop on What's Next in Multimodal Foundational Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02307 2025-08-05 cs.CV cs.LG 57%

Whole-body Representation Learning For Competing Preclinical Disease Risk Assessment

Dmitrii Seletkov, Sophie Starck, Ayhan Can Erdur, Yundi Zhang, Daniel Rueckert, Rickmer Braren

机构 * Institute of Diagnostic Interventional Radiology, Technical University of Munich, School of Medicine, Munich, Germany Chair for AI in Healthcare Medicine, Technical University of Munich (TUM) TUM University Hospital, Munich, Germany Department of Computing, Imperial College London, London, UK German Cancer Consortium (DKTK), Munich partner site, Heidelberg, Germany

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01704 2025-08-05 cs.CV 57%

LT-Gaussian: Long-Term Map Update Using 3D Gaussian Splatting for Autonomous Driving

Luqi Cheng, Zhangshuo Qi, Zijie Zhou, Chao Lu, Guangming Xiong

机构 * Beijing Institute of Technology(北京理工大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments Accepted by IV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01016 2025-08-05 eess.IV cs.CV 57%

Diagnostic Accuracy of Open-Source Vision-Language Models on Diverse Medical Imaging Tasks

Gustav Müller-Franzes, Debora Jutz, Jakob Nikolas Kather, Christiane Kuhl, Sven Nebelung, Daniel Truhn

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18277 2025-08-05 cs.CV cs.LG 57%

Towards Modality Generalization: A Benchmark and Prospective Analysis

Xiaohao Liu, Xiaobo Xia, Zhuo Huang, See-Kiong Ng, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) The University of Sydney(悉尼大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments Accepted by ACM MM 2025 (CR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02342 2025-08-05 cs.IR 50%

Agentic Personalized Fashion Recommendation in the Age of Generative AI: Challenges, Opportunities, and Evaluation

Yashar Deldjoo, Nima Rafiee, Mahdyar Ravanbakhsh

专题命中 多模态评测 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11884 2025-08-05 cs.LG 50%

Out-of-Distribution Detection: A Task-Oriented Survey of Recent Advances

Shuo Lu, Yingsheng Wang, Lijun Sheng, Lingxiao He, Aihua Zheng, Jian Liang

机构 * NLPR \& MAIS, Institute of Automation, Chinese Academy of Sciences Anhui University Hefei China University of Science Anhui University

专题命中 多模态评测 :multi-modal(abstract)

Comments Accepted to ACM Computing Surveys (CSUR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏