AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward Optimization
专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
机构 * LMU Munich(慕尼黑大学) ; Technical University of Munich(慕尼黑技术大学) ; Siemens AG(西门子股份公司) ; University of Science and Technology of China(中国科学技术大学) ; Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) ; Konrad Zuse School of Excellence in Reliable AI (relAI)(Konrad Zuse可靠性人工智能卓越学院) ; University of Oxford(牛津大学)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments Accepted to COLM 2025
机构 * College of Computer Science, Sichuan University(四川大学计算机学院) ; Engineering Research Center of Machine Learning and Industry Intelligence, Ministry of Education, Chengdu, China(教育部机器学习与产业智能工程研究中心) ; Deep NeuroCognition Lab, I2R and CFAR, Agency for Science, Technology and Research, Singapore(深度神经认知实验室,I2R和CFAR,科技研究局,新加坡)
专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV
Comments 10 pages, 7figures
机构 * Department of Engineering Science, University of Oxford(牛津大学工程科学系) ; Department of Electrical Engineering, Imperial College London(伦敦帝国学院电子工程系)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL
Comments 8 pages, 4 figures
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments Accepted by TPAMI2025
Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, Jul. 2025, pp. 5672-5689, vol. 47
机构 * Google DeepMind(谷歌DeepMind)
专题命中 图文多模态 :image-text(abstract);分类 cs.AI
Comments COLM 2025