AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward Optimization
专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
机构 * LMU Munich(慕尼黑大学) ; Technical University of Munich(慕尼黑技术大学) ; Siemens AG(西门子股份公司) ; University of Science and Technology of China(中国科学技术大学) ; Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) ; Konrad Zuse School of Excellence in Reliable AI (relAI)(Konrad Zuse可靠性人工智能卓越学院) ; University of Oxford(牛津大学)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments Accepted to COLM 2025
机构 * College of Computer Science, Sichuan University(四川大学计算机学院) ; Engineering Research Center of Machine Learning and Industry Intelligence, Ministry of Education, Chengdu, China(教育部机器学习与产业智能工程研究中心) ; Deep NeuroCognition Lab, I2R and CFAR, Agency for Science, Technology and Research, Singapore(深度神经认知实验室,I2R和CFAR,科技研究局,新加坡)
专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV
Comments 10 pages, 7figures
机构 * Department of Engineering Science, University of Oxford(牛津大学工程科学系) ; Department of Electrical Engineering, Imperial College London(伦敦帝国学院电子工程系)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL
Comments 8 pages, 4 figures
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments Accepted by TPAMI2025
Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, Jul. 2025, pp. 5672-5689, vol. 47
机构 * Google DeepMind(谷歌DeepMind)
专题命中 图文多模态 :image-text(abstract);分类 cs.AI
Comments COLM 2025
专题命中 音频语音多模态 :cross-modal(title,abstract);audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
机构 * Imperial College London(伦敦帝国学院)
专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(title,abstract);分类 cs.CV、cs.MM、eess.AS
Comments Accepted to IEEE ASRU 2025
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM、eess.AS
Comments Project page: https://github.com/jasongief/TGS-Agent
机构 * Great Bay University(大亚湾大学) ; Nanyang Technological University(南洋理工大学)
专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV
Comments Accepted by Workshop SVC of ACM MM 2025
机构 * Department of Computer Science, The University of Texas at Dallas(计算机科学系,德克萨斯大学达拉斯分校)
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV
机构 * AImotion Bavaria Technische Hochschule Ingolstadt(巴伐利亚AImotion技术大学)
专题命中 音频语音多模态 :multi-modal(title,abstract);分类 cs.AI、cs.MM
Comments 5 pages and 4 figures
机构 * DISI, University of Bologna(DISI,博洛尼亚大学)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments 12th Workshop on Argument Mining (ArgMining 2025) @ ACL 2025
Journal ref In Proceedings of the 12th Argument mining Workshop (ArgMining 2025), pages 388-397, Vienna, Austria
机构 * Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) ; Sony Group Corporation(索尼集团) ; Sony AI, Sony Group Corporation(索尼人工智能,索尼集团)
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 cs.CV、eess.AS
专题命中 音频语音多模态 :audio-visual(title,abstract);分类 eess.AS
Comments 9 pages
专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.MM、eess.AS
Comments Project page: https://ali-vosoughi.github.io/SoundCLIP/
机构 * Brown University(布朗大学) ; University of California at Berkeley(加州大学伯克利分校) ; University of Central Florida(中央佛罗里达大学) ; Dolby Laboratories(杜比实验室)
专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 eess.AS
Comments Updated abstract
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments 3 pages, 2 tables, submitted for arXiv preprint
机构 * Lenovo Research(联想研究院) ; School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院) ; Institute of Automation, CAS(中国科学院自动化研究所) ; Macau University of Science and Technology(澳门科学理工学院) ; School of Computer Science and Technology, UCAS(中国科学院大学计算机科学与技术学院)
专题命中 视频多模态 :multimodal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
Comments preprint
机构 * School of Computer Science, University of Sheffield, UK(计算机科学学院,谢菲尔德大学)
专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV
Journal ref ICCV 2025
机构 * Institute of Digital Games University of Malta(数字游戏研究所马耳他大学) ; Department of Artificial Intelligence University of Malta(人工智能系马耳他大学)
专题命中 跨模态检索 :multimodal(title);分类 cs.MM
专题命中 跨模态检索 :image-text(abstract);分类 cs.CV
机构 * Institute of Trustworthy Embodied AI(可信具身人工智能研究院) ; Fudan University(复旦大学) ; Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心) ; Bytedance Intelligent Creation(字节跳动智能创作)
专题命中 多模态生成 :multimodal(title,abstract);image-text(abstract);分类 cs.CV
Comments Accepted by ICCV 2025
机构 * independent researcher(独立研究者)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI
机构 * University at Buffalo–SUNY(布法罗大学-纽约州立大学) ; Department of Electrical Engineering(电气工程系) ; New Jersey Institute of Technology (NJIT)(新泽西理工学院)
专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI、cs.MM
Comments 16 pages, 4 Figures, 8 Tables
机构 * Guangdong Provincial Key Laboratory of Fully Actuated System Control Theory and Technology, the Southern University of Science and Technology, Shenzhen 518055, China(广东省全自动化系统控制理论与技术重点实验室,南方科技大学,深圳518055,中国) ; Department of Mechanical Engineering, City University of Hong Kong, Hong Kong SAR, China(香港城市大学机械工程系,香港特别行政区,中国)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
Comments To be presented at the 28th IEEE International Conference on Intelligent Transportation Systems (ITSC), 2025
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Shanghai Jiao Tong University(上海交通大学) ; The Chinese University of Hong Kong(香港中文大学) ; Beihang University(北京航空航天大学)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV
专题命中 多模态生成 :multi-modal(title,abstract)
Comments Accepted at IEEE International Conference on Distributed Computing Systems (ICDCS 2025)
机构 * Fudan University(复旦大学) ; Tencent, YouTu Lab(腾讯、YouTu实验室)
专题命中 多模态生成 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV
Comments Accepted by ACM MM'25. arXiv admin note: text overlap with arXiv:2409.03270
机构 * The Hong Kong Polytechnic University(香港理工大学)
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI
Comments 10 pages, 3 figures, and published to ICCV2025