LinMU: Multimodal Understanding Made Linear
LinMU: 使多模态理解线性化
机构 * Princeton University(普林斯顿大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM
AI总结 LinMU通过线性复杂度设计实现多模态理解,无需二次注意力模块,提升视频处理效率。
Comments 23 pages, 7 figures
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
LinMU: 使多模态理解线性化
机构 * Princeton University(普林斯顿大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM
AI总结 LinMU通过线性复杂度设计实现多模态理解,无需二次注意力模块,提升视频处理效率。
Comments 23 pages, 7 figures
通过GPU内部调度和资源共享实现解耦的多阶段MLLM推理
机构 * Wuhan University(武汉大学)
专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract)
AI总结 通过GPU内部调度和资源共享实现解耦的多阶段MLLM推理,提升吞吐量和延迟性能
多模态机器学习在人机交互中早期信任预测中的应用:利用面部图像和GSR生物信号
专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract)
AI总结 本研究提出多模态机器学习框架,结合面部图像和GSR生物信号,用于预测人机交互中AI或人类推荐的早期信任,通过多模态堆叠集成提升预测性能。
Comments This version contains errors in content presentation and arrangement, so it is being withdrawn until a corrected version is generated
Mavors:多粒度视频表示用于多模态大语言模型
机构 * Peking University(北京大学) ; Kling Team(Kling团队) ; Nanjing University(南京大学) ; CASIA(中国科学院自动化研究所)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 Mavors通过多粒度视频表示方法,提升多模态大语言模型在长视频理解中的时空模式保留与计算效率平衡能力。
Comments 22 pages
机构 * Department of Computer Science and Engineering, Southern University of Science and Technology(计算机科学与工程系,南方科技大学) ; School of Computer Science, University of Birmingham(计算机科学学院,伯明翰大学) ; Department of Computer Science, University of Hong Kong(计算机科学系,香港大学) ; Division of Informatics, Imaging and Data Sciences, University of Manchester(信息学、成像与数据科学系,曼彻斯特大学) ; William and Mary(威廉与玛丽学院) ; Harbin Institute of Technology(哈尔滨工业大学) ; UCAS-Terminus AI Lab, University of Chinese Academy of Sciences(中国科学院大学-Terminus AI实验室)
专题命中 视频多模态 :audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS
Comments Published on IEEE TPAMI
Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 11, pp. 10280-10294, August 2025
专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract)
Comments 11 pages, 6 figures,
机构 * Zhejiang University(浙江大学)
专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract)
机构 * 2 Department of Electrical ; Computer Engineering University of Waterloo, Waterloo, ON, Canada N2L 3G1 Email ; 3 College of Computer ; Information Sciences Prince Sultan University Email
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Journal ref 2024 IEEE International Symposium on Medical Measurements and Applications (MeMeA)
机构 * NTT, Inc.(日本NTT公司)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM
Comments Accepted at APSIPA ASC 2025
机构 * Southern University of Science and Technology(南方科技大学) ; Tencent Inc.(腾讯公司)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted by ACCV 2022
机构 * Dept. Information Engineering, Electronics and Telecommunications (DIET), Sapienza University of Rome(信息工程、电子与电信系(DIET),罗马萨皮恩扎大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS
Comments Acepted at IJCNN 2025
机构 * University Of Southern California(美国南加州大学)
专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Project Page: https://deeptracereward.github.io/
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract)
Comments Accepted by 2025 Recsys EARL Workshop
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)
机构 * Imperial College London(帝国理工学院伦敦分校)
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)
Comments 10 pages, 1 figure
机构 * IoT Research Laboratory, Ontario Tech University(Ontario Tech 大学物联网研究实验室) ; Ontario Shores Centre for Mental Health Sciences(Ontario Shores 精神健康科学中心) ; Temerty Faculty of Medicine, University of Toronto(多伦多大学Temerty医学学院) ; Faculty of Science, Ontario Tech University(Ontario Tech 大学科学学院)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM
机构 * e-Media Research Lab(e-Media研究实验室) ; ESAT-STADIUS Division, KU Leuven(ESAT-STADIUS部门,KU莱顿大学) ; M-Group, DistriNet, Department of Computer Science, KU Leuven(M组、DistriNet、计算机科学系,KU莱顿大学)
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract)
Comments This manuscript has been submitted to a peer-reviewed journal and is currently under review
机构 * Provable Responsible AI and Data Analytics (PRADA) Lab(可证明负责任的人工智能与数据分析实验室) ; King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹大学) ; University of Electronic Science and Technology of China(中国电子科学技术大学) ; University of Copenhagen(哥本哈根大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
机构 * Kuaishou Technology(快手科技) ; Shandong University(山东省大学)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.MM
机构 * Northeastern University(东北大学) ; Shenyang Women’s and Children’s Hospital(沈阳市妇女儿童医院)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract)
Comments 24 pages,12 figures
机构 * Czech Institute of Informatics, Robotics and Cybernetics(捷克信息学、机器人学与自动控制研究所) ; Czech Technical University in Prague(布拉格捷克技术大学)
专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract)
Comments 7 pages, 5 figures, 2 tables, conference
Journal ref 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
机构 * Faculty of Electronic and Information Engineering, Xi’an Jiaotong University(电子与信息工程学院,西安交通大学) ; Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人工智能与机器人研究院,西安交通大学) ; State Key Lab of Brain-Machine Intelligence, Zhejiang University(脑机智能国家重点实验室,浙江大学) ; School of Computer Science, Dalian University of Technology(计算机科学学院,大连理工大学) ; National Key Lab of Human-Machine Hybrid Augmented Intelligence, Xi’an Jiaotong University(人机混合增强智能国家实验室,西安交通大学)
专题命中 视频多模态 :multi-modal(title);multimodal(abstract);cross-modal(abstract)
机构 * Department of Computer Science, University of Exeter(埃克塞特大学计算机科学系) ; Institute for Analytics and Data Science, University of Essex(埃塞克斯大学分析与数据科学研究所)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM
Comments ICDMW 2024, Github: https://github.com/EvelynZ10/cmfusion
Journal ref 2024 IEEE International Conference on Data Mining Workshops (ICDMW), Abu Dhabi, United Arab Emirates, 2024, pp. 183-190
机构 * University of Science and Technology of China(中国科学技术大学) ; Tongyi Lab, Alibaba Group(阿里巴巴集团通义实验室) ; Tsinghua University(清华大学)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to CVPR2025
机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments ICLR 2025; first two authors contributed equally. Project page: https://CREMA-VideoLLM.github.io/
机构 * University of Edinburgh(爱丁堡大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted at the ICLR 2025 Workshop on Reasoning and Planning for Large Language Models
专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract)
Comments Accepted in Proc. ACM Interactive, Mobile, Wearable and Ubiquitous Technologies,(March 2025), 23 pages. https://doi.org/10.1145/3712284
机构 * School of Computer Science ; Technology , Xinjiang University, Urumqi, China Email
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI、eess.AS
Comments Accepted for publication by 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)