Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL
Comments Published in IEEE Access
Journal ref IEEE Access, vol. 13, pp. 176751-176769, 2025
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL
Comments Published in IEEE Access
Journal ref IEEE Access, vol. 13, pp. 176751-176769, 2025
机构 * Advanced Micro Devices, Inc.(Advanced Micro Devices公司)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
机构 * BAAI(百度人工智能研究院)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
Comments project page: https://emu.world
机构 * Department of Geography(地理系) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV
Comments 41 pages
机构 * Acellera Labs(Acellera实验室) ; ICREA, Universitat Pompeu Fabra(ICREA、庞培法布拉大学)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI
机构 * Zhejiang Sci-Tech University ; Style3D Research Hangzhou China ; State Key Lab of CAD\&CG, Zhejiang University ; Shanghai Jiao Tong University Shanghai China ; State Key Lab of CAD\&CG, Zhejiang University Hangzhou China ; Style3D Research ; Shanghai Jiao Tong University
专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV
Comments 23 pages,20 figures
Journal ref ACM Trans. Graph. 44, 6, Article 216 (2025)
机构 * City University of Hong Kong(香港城市大学) ; WeChat, Tencent Inc(微信、腾讯公司) ; Manycore Tech Inc(很多核科技公司)
专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV
Comments Code, model, and dataset will be released at project page soon: https://luckyhzt.github.io/x2video
机构 * CUHK(香港中文大学) ; HKUST(香港科技大学) ; HKU(香港大学) ; ByteDance Inc(字节跳动公司)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
机构 * Shanghai AI Laboratory(上海人工智能实验室) ; Shanghai Innovation Institute(上海创新研究院) ; Nanjing University(南京大学) ; The University of Sydney(悉尼大学) ; Shanghai Jiao Tong University(上海交通大学) ; Tsinghua University(清华大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV
Comments 33 pages, 13 figures, 10 tables
机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学) ; Guangdong University of Technology(广东工业大学) ; The Hong Kong Polytechnic University(香港理工大学) ; The Chinese University of Hong Kong(香港中文大学) ; Guizhou University(贵州大学) ; Zhejiang University(浙江大学) ; Nanyang Technological University(南洋理工大学)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI
机构 * University of Chinese Academy of Sciences(中国科学院大学)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI
Comments Project Page: https://github.com/martian422/MaskGRPO
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL
Comments This paper is accepted for presentation in TRB annual meeting 2026. The version presented here is the preprint version before peer review process
机构 * National University of Singapore(新加坡国立大学) ; Fudan University(复旦大学) ; Harbin Institute of Technology(哈尔滨工业大学) ; Eastern Institute of Technology(东方技术研究所)
专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CL
Comments Accepted by ACM MM 2025
专题命中 多模态生成 :any-to-any(title,abstract);分类 cs.CV
Comments Accepted at ICCV25
机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; Horizon Robotics ; Nanjing University(南京大学) ; Huazhong University of Science & Technology(华中科技大学) ; Shanghai AI Laboratory(上海人工智能实验室)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
机构 * Baidu VIS(百度视觉) ; National University of Singapore(新加坡国立大学)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
Comments 23 pages, 10 figures
专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV
机构 * State Key Laboratory of Media Convergence and Communication(媒体融合与传播国家重点实验室) ; Communication University of China(中国传媒大学) ; Microsoft Research Asia(微软亚洲研究院)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI
机构 * ByteDance Seed(字节跳动种子)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
Comments Technical Report
机构 * Massachusetts Institute of Technology(麻省理工学院)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI
Comments 9 figures, 15 pages. Accepted and soon published in the ASME Journal of Mechanical Design
机构 * LiAuto Inc(LiAuto公司) ; Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) ; School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动研究所)
专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV
Comments Under review
机构 * State Key Laboratory of Robotics and Systems, Harbin Institute of Technology(机器人系统国家重点实验室,哈尔滨工业大学)
专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV
机构 * Peking University(北京大学) ; Microsoft(微软)
专题命中 多模态生成 :MLLM(title);multimodal(abstract);分类 cs.AI
Comments Accepted to EMNLP 2025 Main Conference
机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室) ; ByteDance(字节跳动)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
Comments ICLR 2025
机构 * King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹大学科学与技术大学) ; Lanzhou University(兰州大学) ; Meta AI ; The University of Sydney(悉尼大学) ; IHPC, A*STAR(IHPC,A*STAR)
专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.CV
Comments ICCV 2025, Project in https://wikiautogen.github.io/
机构 * ETH Zürich(苏黎世联邦理工学院) ; DisneyResearch | Studios(迪士尼研究室)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
Comments Added more evaluation since the first version. Accepted to SMI 2025. Computers & Graphics
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
Comments Project Page: https://personavlog-paper.github.io/
机构 * Peking University(北京大学) ; Wuhan University(武汉大学) ; Shanghai University of International Business(上海国际商务大学) ; Microsoft(微软)
专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI
Comments SIGIR 2025
Journal ref SIGIR 2025: Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval Pages 581 - 591
机构 * MAUM AI Inc.(MAUM AI公司) ; Artificial Intelligence Graduate School UNIST(UNIST人工智能研究生学院)
专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV
Comments Accepted to BMVC 2025
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI
Journal ref Biomed. Eng. Lett. 15 (2025)