arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-09-29 至 2025-09-29 共收录 12 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 12 篇

2509.21360 2025-09-29 cs.CV cs.AI 81%

Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models

Xingkai Peng, Jun Jiang, Meng Tong, Shuai Li, Weiming Zhang, Nenghai Yu, Kejiang Chen

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22570 2025-09-29 cs.AI 79%

UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration

Qi Mao, Tinghan Yang, Jiahao Li, Bin Li, Libiao Jin, Yan Lu

机构 * State Key Laboratory of Media Convergence and Communication(媒体融合与传播国家重点实验室) Communication University of China(中国传媒大学) Microsoft Research Asia(微软亚洲研究院)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21874 2025-09-29 cs.LG 78%

Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models

Yifei Peng, Yaoli Liu, Enbo Xia, Yu Jin, Wang-Zhou Dai, Zhong Ren, Yao-Xiang Ding, Kun Zhou

机构 * State Key Laboratory of CAD&CG(计算机辅助设计与图形学国家重点实验室) National Key Laboratory for Novel Software Technology(新型软件技术国家实验室)

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03255 2025-09-29 cs.CV 70%

DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation

Qingdong He, Jinlong Peng, Pengcheng Xu, Boyuan Jiang, Xiaobin Hu, Donghao Luo, Yong Liu, Yabiao Wang, Chengjie Wang, Xiangtai Li, Jiangning Zhang

机构 * Youtu Lab, Tencent(腾讯优图实验室) Western University(西部大学) Nanyang Technological University(南洋理工大学) Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21887 2025-09-29 cs.CV cs.MM 62%

StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing

Liyang Chen, Tianze Zhou, Xu He, Boshi Tang, Zhiyong Wu, Yang Huang, Yang Wu, Zhongqian Sun, Wei Yang, Helen Meng

专题命中 多模态生成 :audio-visual(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21375 2025-09-29 cs.CV cs.AI 62%

Automated Prompt Generation for Creative and Counterfactual Text-to-image Synthesis

Aleksa Jelaca, Ying Jiao, Chang Tian, Marie-Francine Moens

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments text-to-image generation, automatic prompt, DPO, Counterfactual

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22573 2025-09-29 cs.RO cs.CV 57%

MINT-RVAE: Multi-Cues Intention Prediction of Human-Robot Interaction using Human Pose and Emotion Information from RGB-only Camera Data

Farida Mohsen, Ali Safa

机构 * College of Science and Engineering, Hamad Bin Khalifa University(哈马德·本·哈利法大学科学与工程学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22485 2025-09-29 cs.CV 57%

Group Critical-token Policy Optimization for Autoregressive Image Generation

Guohui Zhang, Hu Yu, Xiaoxiao Ma, JingHao Zhang, Yaning Pan, Mingde Yao, Jie Xiao, Linjiang Huang, Feng Zhao

机构 * University of Science and Technology of China(中国科学技术大学) Shanghai Innovation Institute(上海创新研究院) Fudan University(复旦大学) CUHK(香港中文大学) Beihang University(北京航空航天大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

Comments Code is available at https://github.com/zghhui/GCPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21997 2025-09-29 cs.CV 57%

Exposing Hallucinations To Suppress Them: VLMs Representation Editing With Generative Anchors

Youxu Shi, Suorong Yang, Dong Liu

机构 * University of Science and Technology of China(中国科学技术大学) Nanjing University(南京大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21760 2025-09-29 cs.CV 57%

UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models

Lan Chen, Yuchao Gu, Qi Mao

机构 * MIPG, Communication University of China(信息与通信大学) Show Lab, National University of Singapore(新加坡国立大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20253 2025-09-29 cs.RO cs.AI 57%

AnchDrive: Bootstrapping Diffusion Policies with Hybrid Trajectory Anchors for End-to-End Driving

Jinhao Chai, Anqing Jiang, Hao Jiang, Shiyi Mu, Zichong Gu, Hao Sun, Shugong Xu

机构 * School of Communication and Information Engineering, Shanghai University, Shanghai 200444, China(信息工程学院,上海大学) Bosch Corporate Research, Bosch (China) Investment Ltd., Shanghai, China(博世企业研究,博世(中国)投资有限公司) School of Mechanical Engineering, Shanghai Jiao Tong University, Shanghai, China(机械工程学院,上海交通大学) Xi'an Jiaotong-Liverpool University, Suzhou, China(西安交通大学利物浦大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21664 2025-09-29 cs.RO cs.LG 50%

Generating Stable Placements via Physics-guided Diffusion Models

Philippe Nadeau, Miguel Rogel, Ivan Bilić, Ivan Petrović, Jonathan Kelly

机构 * STARS Laboratory, University of Toronto Institute for Aerospace Studies(多伦多大学航空航天研究所STARS实验室) Laboratory for Autonomous Systems and Mobile Robotics(自主系统与移动机器人实验室) University of Zagreb Faculty of Electrical Engineering and Computing(Zagreb大学电气工程与计算学院)

专题命中 多模态生成 :multi-modal(abstract)

Comments Submitted to the IEEE International Conference on Robotics and Automation 2026, Vienna, Austria, June 1-5, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏