Dreaming User Multimodal Representation Guided by The Platonic Representation Hypothesis for Micro-Video Recommendation
专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI
Comments 4 Figure; 2 Table
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI
Comments 4 Figure; 2 Table
专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.AI
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM
Comments Accepted at ICME 2024 as poster presentation. arXiv admin note: text overlap with arXiv:2306.14392
专题命中 视频多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI
Comments Preprint. Work in progress
专题命中 视频多模态 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM
专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.CL
Comments Accepted by IEEE Access
Journal ref in IEEE Access, vol. 11, pp. 51229-51240, 2023
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI
Comments Published in Transactions on Machine Learning Research. Previous version accepted to Foundation Models for Decision Making Workshop at NeurIPS 2022
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.MM
Comments Accepted to TPAMI 2023
专题命中 视频多模态 :multi-modal(title);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
Comments Accepted to CVPR 2023
专题命中 视频多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.CL
Journal ref 2023 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW)
专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM
Comments 12 pages; accepted by IEEE TMM
Journal ref IEEE Transactions on Multimedia 2021
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL
Comments EACL 2023
专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI
专题命中 视频多模态 :cross-modal(title);multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.MM
Comments ACM MM 2022
专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL
Comments Accepted by Conference on Empirical Methods in Natural Language Processing (EMNLP 2021, Oral)
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.MM
Comments Accepted by ACL-2022
专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM
Comments Accepted by TPAMI 2021
专题命中 视频多模态 :multi-modal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.CL
Comments Accepted at AAAI 2021
专题命中 视频多模态 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL
Comments Accepted by ICASSP 2020
专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM
Comments 22 pages, submitted to ACM Transactions on Multimedia Computing Communications and Applications(ACM TOMM)
专题命中 视频多模态 :multi-modal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.CL
Comments Accepted for an Oral presentation at the DSTC7 workshop at AAAI 2019
专题命中 视频多模态 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL
Comments To appear in AAAI 2016
专题命中 视频多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.MM
Comments 7 pages
Journal ref Journal of Intelligent Computing, Volume: 1, Issue: 4 (December 2010), Page: 165-175
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI;multimodal foundation model(comments)
Comments Accepted at "What is Next in Multimodal Foundation Models?" (MMFM) workshop at ICCV 2023
Zero-MELO:基于多模态大语言模型的测试时证据校准用于零样本微手势识别
机构 * University of Oulu(奥卢大学) ; Peking University(北京大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
AI总结 该研究针对多模态大语言模型在微手势识别中局部证据不足、分数偏差的瓶颈,提出Zero-MELO框架,结合树搜索、测试时校准与多线索融合,在iMiGUE和MA-52数据集上显著优于Qwen2.5-VL基线。
Comments Accepted by ACM MM 2026
Med-CRAFT:通过知识图谱遍历自动构建可解释的多跳视频工作负载
机构 * Beijing Institute of Technology(北京理工大学) ; The Hong Kong Polytechnic University(香港理工大学) ; Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract_cn);分类 cs.AI
AI总结 Med-CRAFT通过知识图谱遍历自动构建可解释的多跳视频工作负载,生成具有细粒度时间选择性和多跳逻辑复杂性的医疗视频推理基准。
Comments 23 pages, 4 figures, 10 tables
用于多模态纵向图像插补与插值的隐式神经表示
机构 * University Hospital Augsburg(奥格斯堡大学医院) ; University of Augsburg(奥格斯堡大学) ; Technical University of Munich(慕尼黑工业大学) ; Bavarian Cancer Research Center (BZKF)(巴伐利亚癌症研究中心) ; Swabian Children’s Cancer Center(施瓦本儿童癌症中心)
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
AI总结 该研究针对临床纵向MRI数据缺失等问题,提出条件隐式神经表示模型,经儿科脑肿瘤数据验证,可显著改进插值效果,置信度估计可靠。
Farm-LightSeek:一种以边缘为中心的多模态农业物联网数据分析框架,集成轻量级语言模型
机构 * School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院) ; School of Engineering, Swinburne University of Technology(斯威本科技大学工程学院) ; School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) ; School of Computing, Engineering and Mathematical Sciences, La Trobe University(拉筹伯大学计算、工程与数学科学学院)
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
AI总结 针对智能农业面临的挑战,提出Farm-LightSeek框架,将大语言模型与边缘计算集成。通过传感器收集多源数据,在边缘节点进行跨模态推理等。创新包括闭环架构等,实验表明该框架在关键任务中性能可靠,推动了智能实时农业及两者深度集成。
Comments Accepted by IEEE Internet of Things Magazine
超越独立优化:多模态边缘智能中的压缩、混合专家路由和量化交互
机构 * Nirma University(尼玛大学) ; Singapore Institute of Technology(新加坡理工学院) ; Marwadi University(马尔瓦迪大学)
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI
AI总结 研究多模态边缘智能中高效推理受多种因素限制,回顾相关模型进展,指出技术间相互影响不能独立优化,介绍关键设计权衡,引入视频MoE模型诊断方法,强调多方面开放研究方向。
学习检测跨模态否定:潜在表示分析与基于注意力的解决方案
专题命中 视频多模态 :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI
AI总结 研究跨模态否定检测难题,提出新型跨模态注意力架构,分析发现文本与视觉否定的不对称性,结合自监督视频表示推进时间否定建模,为多模态系统学习语义对齐表示提供新方法。
Comments This manuscript is an accepted version of the article (published at ICNLP2026). Published in IEEE Xplore, DOI:10.1109/ICNLP69856.2026.11527861 document: https://ieeexplore.ieee.org/abstract/document/11527861
Journal ref 2026 8th International Conference on Natural Language Processing (ICNLP), Xi'an, China, 2026, pp. 613-622