ThinkFake: Reasoning in Multimodal Large Language Models for AI-Generated Image Detection
专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV
机构 * Shanghai AI Lab(上海人工智能实验室)
专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV
机构 * Adobe Research(Adobe研究院) ; University of Maryland(马里兰大学) ; University at Buffalo(布法罗大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
机构 * Indian Institute of Technology Patna(印度帕纳布理工大学) ; Microsoft, India(微软印度)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI
机构 * Chinese University of Hong Kong(中国香港大学) ; City University of Hong Kong(香港城市大学) ; University of Oxford(牛津大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; The Sixth Affiliated Hospital, Sun Yat-sen University(中山大学第六附属医院)
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV
Comments 40 pages, 22 figures; Accepted by NeurIPS 2025 Dataset and Benchmark Track
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
Comments Project page with our code and dataset: https://irvlutd.github.io/MultiGrounding
机构 * University of Pennsylvania(宾夕法尼亚大学) ; BigHat Biosciences(BigHat生物技术公司)
专题命中 多模态评测 :multimodal(title,abstract);multi-modal(comments)
Comments NeurIPS 2025 AI4Science Workshop and NeurIPS 2025 Multi-modal Foundation Models and Large Language Models for Life Sciences Workshop
机构 * Unstructured Technologies
专题命中 多模态评测 :multi-modal(abstract);分类 cs.CL、cs.AI
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Microsoft(微软) ; University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校) ; Amazon(亚马逊) ; Amazon AGI(亚马逊人工智能实验室)
专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI
Comments major revision; few examples of changes: added contemporary LLMs and new SOTA model, improved readability, expanded related work, etc
机构 * Nankai Institute of Advanced Research (SHENZHEN FUTIAN)(南开先进研究院(深圳福田)) ; College of Computer Science & VCIP, Nankai University(计算机科学与VCIP学院,南开大学) ; School of Computing, Australian National University(计算学院,澳大利亚国立大学) ; Graduate School of Science and Technology, Keio University(科学与技术研究生院,庆应大学) ; Department of Electronic Engineering, Tsinghua University(电子工程系,清华大学) ; Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
专题命中 多模态评测 :multimodal(abstract,comments);分类 cs.CV
Comments [Work in progress] A comprehensive survey of intelligent colonoscopy in the multimodal era. [Updated Version V2] New training strategy for colonoscopy-specific multimodal language model
Journal ref Machine Intelligence Research 2025
机构 * Georgia Institute of Technology(佐治亚理工学院)
专题命中 多模态评测 :multimodal(abstract);分类 cs.CL
Comments Accepted to EMNLP 2025 Main Conference
专题命中 多模态评测 :multimodal(abstract)
机构 * Hamad Bin Khalifa University(哈马德·本·卡尔法大学) ; Maastricht University(马斯特里赫特大学) ; KTH Royal Institute of Technology(皇家理工学院)
专题命中 多模态评测 :multimodal(abstract)
机构 * Shenzhen Research Institute of Big Data, School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, Guangdong(深圳大数据研究院,科学与工程学院,香港中文大学(深圳)) ; Shenzhen Research Institute of Big Data, School of Data Science, The Chinese University of Hong Kong, Shenzhen, Guangdong(深圳大数据研究院,数据科学学院,香港中文大学(深圳)) ; Networking and User Experience Lab, Huawei Technologies(网络与用户体验实验室,华为技术)
专题命中 多模态评测 :multi-modal(abstract)