arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-11-11 至 2025-11-11 共收录 8 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 8 篇

2511.07290 2025-11-11 eess.IV cs.CV cs.MM 81%

CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video

Xinyi Wang, Angeliki Katsenou, Junxiao Shen, David Bull

机构 * School of Computer Science, University of Bristol(布里斯托大学计算机科学学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24838 2025-11-11 cs.CV cs.AI 73%

VideoCAD: A Dataset and Model for Learning Long-Horizon 3D CAD UI Interactions from Video

Brandon Man, Ghadi Nehme, Md Ferdous Alam, Faez Ahmed

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 视频多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18094 2025-11-11 cs.CV cs.AI 62%

UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning

Ye Liu, Zongyang Ma, Junfu Pu, Zhongang Qi, Yang Wu, Ying Shan, Chang Wen Chen

机构 * The Hong Kong Polytechnic University(香港理工大学) ARC Lab, Tencent PCG(腾讯PCG ARC实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) vivo Mobile Communication Co.(vivo移动通信公司) MindWingman Technology (Shenzhen) Co., Ltd.(深圳MindWingman技术有限公司)

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2025 Camera Ready. Project Page: https://polyu-chenlab.github.io/unipixel/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06344 2025-11-11 cs.CL cs.AI 62%

TimeSense:Making Large Language Models Proficient in Time-Series Analysis

Zhirui Zhang, Changhua Pei, Tianyi Gao, Zhe Xie, Yibo Hao, Zhaoyang Yu, Longlong Xu, Tong Xiao, Jing Han, Dan Pei

机构 * Tsinghua University(清华大学) Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心) ZTE Corporation(中兴通讯)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07106 2025-11-11 cs.CV 57%

HENet++: Hybrid Encoding and Multi-task Learning for 3D Perception and End-to-end Autonomous Driving

Zhongyu Xia, Zhiwei Lin, Yongtao Wang, Ming-Hsuan Yang

机构 * University of California, Merced(加州大学默塞德分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

Comments Preliminary version, 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07049 2025-11-11 cs.CV cs.CR 57%

From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge

Hui Lu, Yi Yu, Song Xia, Yiming Yang, Deepu Rajan, Boon Poh Ng, Alex Kot, Xudong Jiang

专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV

Comments AAAI 2026 (Oral presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06281 2025-11-11 cs.CV 57%

VideoSSR: Video Self-Supervised Reinforcement Learning

Zefeng He, Xiaoye Qu, Yafu Li, Siyuan Huang, Daizong Liu, Yu Cheng

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学) Wuhan University(武汉大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05555 2025-11-11 cs.CY cs.SI 50%

Deception Decoder: Proposing a Human-Focused Framework for Identifying AI-Generated Content on Social Media

C. Bowman Kerbage

专题命中 视频多模态 :multimodal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏