arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-12 至 2025-09-12 共收录 13 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 13 篇

2509.09160 2025-09-12 cs.CL cs.AI 84%

Target-oriented Multimodal Sentiment Classification with Counterfactual-enhanced Debiasing

Zhiyue Liu, Fanrong Ma, Xin Ling

机构 * School of Computer, Electronics and Information(计算机、电子与信息学院) Guangxi University(广西大学) Guangxi Key Laboratory of Multimedia Communications(广西多媒体通信与网络技术重点实验室) School of Sociology and Anthropology(社会学与人类学学院) Sun Yat-sen University(中山大学)

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CL、cs.AI

Comments Accepted by the IEEE International Conference on Multimedia and Expo (ICME 2025). © 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09307 2025-09-12 cs.CV cs.AI cs.CL cs.MM 83%

Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization

Zhengzhao Lai, Youbin Zheng, Zhenyang Cai, Haonan Lyu, Jinpu Yang, Hongqing Liang, Yan Hu, Benyou Wang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07084 2025-09-12 cs.RO 82%

DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models

Shucheng Huang, Freda Shi, Chen Sun, Jiaming Zhong, Minghao Ning, Yufeng Yang, Yukun Lu, Hong Wang, Amir Khajepour

机构 * MVSLab, Department of Mechanical and Mechatronics Engineering, University of Waterloo(滑铁卢大学机械与机电工程系MVSLab) CompLING Lab, David R. Cheriton School of Computer Science, University of Waterloo(滑铁卢大学大卫·R·切里顿计算机科学学院CompLING Lab) Department of Data and Systems Engineering, University of Hong Kong(香港大学数据与系统工程系) Department of Mechanical Engineering, University of New Brunswick(新不伦瑞克大学机械工程系) School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动性学院)

专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract)

Comments This work has been accepted to IEEE Transactions on Vehicular Technology. Please refer to the copyright notice for additional information

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09254 2025-09-12 cs.CV cs.MM 81%

Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis

Jing Hao, Yuxuan Fan, Yanpeng Sun, Kaixin Guo, Lizhuo Lin, Jinrong Yang, Qi Yong H. Ai, Lun M. Wong, Hao Tang, Kuo Feng Hung

机构 * Faculty of Dentistry, The University of Hong Kong(香港大学牙科学院) The Hong Kong University of Science and Technology (GZ)(香港科学与技术大学) National University of Singapore(新加坡国立大学) CVTE Sun Yat-sen University(孙中山大学) Department of Diagnostic Radiology, The University of Hong Kong(香港大学放射科) Imaging and Interventional Radiology, Faculty of Medicine, The Chinese University of Hong Kong(香港中文大学医学院影像与介入放射科) School of Computer Science, Peking University(北京大学计算机科学系)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 40 pages, 26 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09014 2025-09-12 cs.CV cs.CL 81%

COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation

Umair Hassan

机构 * Independent Researcher(独立研究者)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments 17 pages, 3 figures, 3 tables. Dataset available at https://huggingface.co/datasets/umairhassan02/urdu-translated-coco-captions-subset. Scripts and notebooks to reproduce results available at https://github.com/umair-hassan2/COCO-Urdu

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09190 2025-09-12 cs.CV 79%

VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models: Methods and Results

Hanwei Zhu, Haoning Wu, Zicheng Zhang, Lingyu Zhu, Yixuan Li, Peilin Chen, Shiqi Wang, Chris Wei Zhou, Linhan Cao, Wei Sun, Xiangyang Zhu, Weixia Zhang, Yucheng Zhu, Jing Liu, Dandan Zhu, Guangtao Zhai, Xiongkuo Min, Zhichao Zhang, Xinyue Li, Shubo Xu, Anh Dao, Yifan Li, Hongyuan Yu, Jiaojiao Yi, Yiding Tian, Yupeng Wu, Feiran Sun, Lijuan Liao, Song Jiang

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments ICCV VQualA Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19662 2025-09-12 physics.ed-ph 78%

Multimodal large language models and physics visual tasks: comparative analysis of performance and costs

Giulia Polverini, Bor Gregorcic

专题命中 多模态评测 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04077 2025-09-12 cs.CL cs.SD eess.AS 73%

A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion Expressions

Chung-Chun Wang, Jhen-Ke Lin, Hao-Chien Lu, Hong-Yun Lin, Berlin Chen

机构 * National Taiwan Normal University(台湾国立台湾师范大学)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CL、eess.AS

Comments submitted to the ISCA SLaTE-2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09473 2025-09-12 cs.CL 57%

Mitigating Language Barriers in Education: Developing Multilingual Digital Learning Materials with Machine Translation

Lucie Poláková, Martin Popel, Věra Kloudová, Michal Novák, Mariia Anisimova, Jiří Balhar

机构 * Charles University, Faculty of Mathematics and Physics(查理大学数学与物理系)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL

Comments 8 pages, 2 figures

Journal ref L. Poláková, M. Popel, V. Kloudová, M. Novák, M. Anisimova, J. Balhar (2025). Mitigating Language Barriers in Education: Developing Multilingual Digital Learning Materials with Machine Translation, EDULEARN25, pp. 8754-8760

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09324 2025-09-12 cs.CV 57%

Fine-Grained Customized Fashion Design with Image-into-Prompt benchmark and dataset from LMM

Hui Li, Yi You, Qiqi Chen, Bingfeng Zhang, George Q. Huang

机构 * The Hong Kong Polytechnic University, China(香港理工大学) China University of Petroleum (East China), China(中国石油大学(华东))

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09227 2025-09-12 eess.IV cs.CV 57%

Dynamic Structural Recovery Parameters Enhance Prediction of Visual Outcomes After Macular Hole Surgery

Yinzheng Zhao, Zhihao Zhao, Rundong Jiang, Louisa Sackewitz, Quanmin Liang, Mathias Maier, Daniel Zapp, Peter Charbel Issa, Mohammad Ali Nasseri

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

Comments TVST

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10546 2025-09-12 cs.CV cs.RO 57%

The Oxford Spires Dataset: Benchmarking Large-Scale LiDAR-Visual Localisation, Reconstruction and Radiance Field Methods

Yifu Tao, Miguel Ángel Muñoz-Bañón, Lintong Zhang, Jiahao Wang, Lanke Frank Tarimo Fu, Maurice Fallon

机构 * Oxford Robotics Inst., Dept. of Eng. Science, Univ. of Oxford, UK(牛津大学机器人研究所、工程科学系) Group of Automation, Robotics and Computer Vision (AUROVA), University of Alicante, Spain(自动化、机器人与计算机视觉小组(AUROVA)、阿尔基兰特大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

Comments Accepted by IJRR. Website: https://dynamic.robots.ox.ac.uk/datasets/oxford-spires/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05019 2025-09-12 cs.CE 50%

FinMultiTime: A Four-Modal Bilingual Dataset for Financial Time-Series Analysis

Wenyan Xu, Dawei Xiang, Yue Liu, Xiyu Wang, Yanxiang Ma, Liang Zhang, Shu Hu, Chang Xu, Jiaheng Zhang

专题命中 多模态评测 :multimodal(abstract)

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏