arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-08-12 至 2025-08-12 共收录 100 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 16 篇

2508.05658 2025-08-12 cs.CR cs.CV cs.MM 81%

Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model Safeguards

Song Yan, Hui Wei, Jinlong Fei, Guoliang Yang, Zhengyu Zhao, Zheng Wang

机构 * Information Engineering University Zhengzhou China School of Computer Science, \ University Wuhan China Xi’an Jiaotong University Xi’an China Wuhan University Wuhan China Information Engineering University School of Computer Science, \ University Xi’an Jiaotong University Wuhan University

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments This paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07519 2025-08-12 cs.CV 79%

Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing

Joonghyuk Shin, Alchan Hwang, Yujin Kim, Daneul Kim, Jaesik Park

机构 * Seoul National University(首尔国立大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

Comments ICCV 2025. Project webpage: https://joonghyuk.com/exploring-mmdit-web/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23186 2025-08-12 cs.CV 79%

HiGarment: Cross-modal Harmony Based Diffusion Model for Flat Sketch to Realistic Garment Image

Junyi Guo, Jingxuan Zhang, Fangyu Wu, Huanda Lu, Qiufeng Wang, Wenmian Yang, Eng Gee Lim, Dongming Lu

机构 * Xi’an Jiaotong Liverpool University(西安交通大学利物浦大学) NingboTech University(宁波科技学院) Beijing Normal University(北京师范大学) Zhejiang University(浙江大学)

专题命中 多模态生成 :cross-modal(title);multi-modal(abstract);分类 cs.CV

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07021 2025-08-12 cs.CV 79%

DocRefine: An Intelligent Framework for Scientific Document Understanding and Content Optimization based on Multimodal Large Model Agents

Kun Qian, Wenjie Li, Tianyu Sun, Wenhong Wang, Wenhan Luo

机构 * Shangqiu University(商丘大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06888 2025-08-12 cs.SE 78%

Multi-Modal Requirements Data-based Acceptance Criteria Generation using LLMs

Fanyu Wang, Chetan Arora, Yonghui Liu, Kaicheng Huang, Chakkrit Tantithamthavorn, Aldeida Aleti, Dishan Sambathkumar, David Lo

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05148 2025-08-12 cs.CV cs.AI cs.LG 73%

LoRA.rar: Learning to Merge LoRAs via Hypernetworks for Subject-Style Conditioned Image Generation

Donald Shenaj, Ondrej Bohdal, Mete Ozay, Pietro Zanuttigh, Umberto Michieli

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments ICCV 2025. Project page: https://donaldssh.github.io/LoRA.rar

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22792 2025-08-12 cs.CV 70%

Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization

Yuxi Zhang, Yueting Li, Xinyu Du, Sibo Wang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10574 2025-08-12 cs.CV cs.MM cs.SD eess.AS 67%

DanceChat: Large Language Model-Guided Music-to-Dance Generation

Qing Wang, Xiaohang Yang, Yilan Dong, Naveen Raj Govindaraj, Gregory Slabaugh, Shanxin Yuan

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07146 2025-08-12 cs.CV cs.AI 62%

Intention-Aware Diffusion Model for Pedestrian Trajectory Prediction

Yu Liu, Zhijie Liu, Xiao Ren, You-Fu Li, He Kong

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08220 2025-08-12 cs.CV 57%

Learning User Preferences for Image Generation Model

Wenyi Mo, Ying Ba, Tianyu Zhang, Yalong Bai, Biye Li

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07540 2025-08-12 cs.CV 57%

CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts

Junuk Cha, Jihyeon Kim

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

Comments ICCVW'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07225 2025-08-12 eess.IV cs.CV q-bio.QM 57%

HaDM-ST: Histology-Assisted Differential Modeling for Spatial Transcriptomics Generation

Xuepeng Liu, Zheng Jiang, Pinan Zhu, Hanyu Liu, Chao Li

机构 * University of Cambridge, UK(剑桥大学,英国) Northeastern University, Shenyang, China(东北大学,中国)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments 10 pages, 5 figures, includes comparisons with TESLA, HiStoGene, and iStar; submitted to arXiv 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16726 2025-08-12 cs.CV cs.LG 57%

EDiT: Efficient Diffusion Transformers with Linear Compressed Attention

Philipp Becker, Abhinav Mehrotra, Ruchika Chavhan, Malcolm Chadwick, Luca Morreale, Mehdi Noroozi, Alberto Gil Ramos, Sourav Bhattacharya

机构 * Samsung, AI Center Cambridge(三星人工智能中心剑桥)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15173 2025-08-12 cs.IR cs.AI cs.ET 57%

Recommendation with Generative Models

Yashar Deldjoo, Zhankui He, Julian McAuley, Anton Korikov, Scott Sanner, Arnau Ramisa, Rene Vidal, Maheswaran Sathiamoorthy, Atoosa Kasrizadeh, Silvia Milano, Francesco Ricci

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

Comments This submission is a full-length book, expanding significantly on two chapters previously submitted (arXiv:2409.10993v1, arXiv:2408.10946v1). It includes additional chapters, context, analysis, and content, providing a comprehensive presentation of the subject. We have ensured it is appropriately presented as a new, distinct work. arXiv admin note: substantial text overlap with arXiv:2409.10993

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10397 2025-08-12 cs.CY 50%

Large Model Empowered Metaverse: State-of-the-Art, Challenges and Opportunities

Yuntao Wang, Qinnan Hu, Zhou Su, Linkang Du, Qichao Xu, Weiwei Li

专题命中 多模态生成 :multimodal(abstract)

Comments 9 pages,5 figures, 1 table, accepted by IEEE Network in Aug. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态评测 24 篇

2508.07766 2025-08-12 cs.CV cs.AI 84%

UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models

Jinke Li, Jiarui Yu, Chenxing Wei, Hande Dong, Qiang Lin, Liangjing Yang, Zhicai Wang, Yanbin Hao

机构 * Zhejiang University(浙江大学) Tencent(腾讯) Shenzhen University(深圳大学) Hefei University of Technology(合肥工业大学)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted at ACM MM 2025 Dataset Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08093 2025-08-12 cs.CV cs.LG cs.MM eess.AS 82%

MDD-Net: Multimodal Depression Detection through Mutual Transformer

Md Rezwanul Haque, Md. Milon Islam, S M Taslim Uddin Raju, Hamdi Altaheri, Lobna Nassar, Fakhri Karray

机构 * Centre for Pattern Analysis and Machine Intelligence, Department of Electrical and Computer Engineering, University of Waterloo(模式分析与机器智能中心,电气与计算机工程系,滑铁卢大学) School of Engineering and Computing, Department of Computer Science and Engineering, American University of Ras Al Khaimah(工程与计算学院,计算机科学与工程系,阿联酋拉线哈姆市美国大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS

Comments Accepted for the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Vienna, Austria

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17632 2025-08-12 cs.AI cs.CV cs.MM 82%

D-Judge: How Far Are We? Assessing the Discrepancies Between AI-synthesized and Natural Images through Multimodal Guidance

Renyang Liu, Ziyu Lyu, Wei Zhou, See-Kiong Ng

机构 * School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学信息科学与技术学院(深圳校区)) Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究所) College of Modern Engineering and the Engineering Research Center of Cyberspace, Yunnan University(云南大学现代工程学院及空天信息工程研究中心)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06272 2025-08-12 cs.CV cs.AI 81%

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance

Zhang Li, Biao Yang, Qiang Liu, Shuo Zhang, Zhiyin Ma, Liang Yin, Linger Deng, Yabo Sun, Yuliang Liu, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07628 2025-08-12 cs.AI 79%

Multimodal AI Systems for Enhanced Laying Hen Welfare Assessment and Productivity Optimization

Daniel Essien, Suresh Neethirajan

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 66 pages, 7 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07596 2025-08-12 cs.CV 79%

From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users

Shahroz Tariq, Simon S. Woo, Priyanka Singh, Irena Irmalasari, Saakshi Gupta, Dev Gupta

机构 * Sungkyunkwan University, S. Korea(顺天大学) University of Queensland, Australia(昆士兰大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

Comments 11 pages, 3 tables, 5 figures, accepted for publicaiton in the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06851 2025-08-12 cs.AI cs.CY 79%

MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams

Pengfei Zhou, Xiaopeng Peng, Fanrui Zhang, Zhaopan Xu, Jiaxin Ai, Yansheng Qiu, Chuanhao Li, Zhen Li, Ming Li, Yukang Feng, Jianwen Sun, Haoquan Zhang, Zizhen Li, Xiaofeng Mao, Zekai Li, Wangbo Zhao, Kai Wang, Xiaojun Chang, Wenqi Shao, Yang You, Kaipeng Zhang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Innovation Institute(上海创新研究院) USTC(中国科学技术大学) RIT(罗切斯特理工学院) HIT(哈尔滨工业大学) WHU(武汉大学) MBZUAI(马克斯·普朗克人工智能研究所) NUS(新加坡国立大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

Comments 35 pages, 33 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06558 2025-08-12 cs.CV cs.LG 79%

On the effectiveness of multimodal privileged knowledge distillation in two vision transformer based diagnostic applications

Simon Baur, Alexandra Benova, Emilio Dolgener Cantú, Jackie Ma

机构 * Fraunhofer Heinrich-Hertz-Institut(弗劳恩霍夫 Heinrich-Hertz 研究所) Universität Osnabrück(奥斯纳布吕克大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16104 2025-08-12 cs.CY 78%

Beauty and the Bias: Exploring the Impact of Attractiveness on Multimodal Large Language Models

Aditya Gulati, Moreno D'Incà, Nicu Sebe, Bruno Lepri, Nuria Oliver

专题命中 多模态评测 :multimodal(title,abstract)

Comments 39 pages, 4 figures, 33 tables; Accepted for publication at the Eighth AAAI/ACM Conference on AI, Ethics and Society (AIES 2025) (https://www.aies-conference.com/2025/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07818 2025-08-12 cs.CV 70%

Segmenting and Understanding: Region-aware Semantic Attention for Fine-grained Image Quality Assessment with Large Language Models

Chenyue Song, Chen Hui, Haiqi Zhu, Feng Jiang, Yachun Mi, Wei Zhang, Shaohui Liu

专题命中 多模态评测 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07554 2025-08-12 cs.MM 70%

FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding

Xusheng He, Wei Liu, Shanshan Ma, Qian Liu, Chenghao Ma, Jianlong Wu

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14533 2025-08-12 cs.CV 70%

ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding

Shuo Cao, Nan Ma, Jiayang Li, Xiaohui Li, Lihao Shao, Kaiwen Zhu, Yu Zhou, Yuandong Pu, Jiarui Wu, Jiaquan Wang, Bo Qu, Wenhai Wang, Yu Qiao, Dajuin Yao, Yihao Liu

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments 43 pages, 31 figures, 13 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06565 2025-08-12 cs.CV cs.LG 70%

Bridging Brain Connectomes and Clinical Reports for Early Alzheimer's Disease Diagnosis

Jing Zhang, Xiaowei Yu, Minheng Chen, Lu Zhang, Tong Chen, Yan Zhuang, Chao Cao, Yanjun Lyu, Li Su, Tianming Liu, Dajiang Zhu

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07022 2025-08-12 cs.AI cs.CL cs.LG cs.MM 67%

MultiMedEdit: A Scenario-Aware Benchmark for Evaluating Knowledge Editing in Medical VQA

Shengtao Wen, Haodong Chen, Yadong Wang, Zhongying Pan, Xiang Chen, Yu Tian, Bo Qian, Dong Liang, Sheng-Jun Huang

专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI、cs.MM

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07714 2025-08-12 cs.CV cs.AI cs.ET 62%

DoorDet: Semi-Automated Multi-Class Door Detection Dataset via Object Detection and Large Language Models

Licheng Zhang, Bach Le, Naveed Akhtar, Tuan Ngo

机构 * The University of Melbourne(墨尔本大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏