Toward Multimodal Image-to-Image Translation
专题命中 其他多模态 :multimodal(title);分类 cs.CV
Comments NIPS 2017 Final paper. v4 updated acknowledgment. Website: https://junyanz.github.io/BicycleGAN/
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 其他多模态 :multimodal(title);分类 cs.CV
Comments NIPS 2017 Final paper. v4 updated acknowledgment. Website: https://junyanz.github.io/BicycleGAN/
专题命中 其他多模态 :multimodal(title);分类 cs.CV
Comments 6 pages, 10 figures
专题命中 其他多模态 :multimodal(title);分类 cs.CV
专题命中 其他多模态 :multimodal(title);分类 cs.CV
Comments Accepted to ISBI 2018
专题命中 其他多模态 :multi-modal(title);分类 cs.CV
Comments S. Andress, A. Johnson, M. Unberath, and A. Winkler have contributed equally and are listed in alphabetical order
Journal ref J. Med. Imag. 5(2), 2018
专题命中 其他多模态 :multi-modal(title);分类 cs.CV
专题命中 其他多模态 :multimodal(title);分类 cs.CL
专题命中 其他多模态 :multimodal(title);分类 cs.CV
Comments This paper has been submitted to Transaction on Medical Imaging
专题命中 其他多模态 :multimodal(title);分类 cs.CV
专题命中 其他多模态 :cross-modal(title);分类 cs.CV
通过概念引导实现鲁棒的上下文分割
机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)
专题命中 其他多模态 :MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI
AI总结 提出概念引导的上下文分割(CG-ICS),通过提取参考图像的高层语义概念而非仅依赖低层视觉匹配,结合文本概念与视觉示例,显著提升分割准确性和鲁棒性。
Comments ECCV 2026
通过连续软化回溯重采样稳定多模态大语言模型的无监督自进化
机构 * Tsinghua University(清华大学) ; Hefei University of Technology(合肥工业大学) ; University of Arizona(亚利桑那大学) ; MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统实验室)
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 本文提出CSRS方法,通过回溯推理机制和软化频率奖励提升多模态大语言模型的无监督自进化稳定性,实验显示在MathVision等基准上表现优异。
Comments 16 pages, 6 figures
BATQuant: 通过可学习的分块优化实现抗异常的MXFP4量化
机构 * Huawei Technologies University of Science(华为技术大学科学)
专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CL、cs.AI
AI总结 本文提出BATQuant方法,通过可学习的分块优化解决MXFP4量化中异常传播问题,实现高性能量化方案。
Comments 30 pages, 13 figures, 7 tables
OCR 或不是?在 MLLMs 时代重新思考文档信息提取:基于真实世界的大规模数据集
机构 * SAP(SAP公司) ; Stanford University(斯坦福大学)
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CL、cs.AI
AI总结 本文探讨了在 MLLMs 时代是否仍需 OCR,通过大规模数据集评估发现,仅图像输入可达到与 OCR 增强方法相当的性能,并展示了通过精心设计的模式、示例和指示可进一步提升 MLLMs 的表现。
MaS-VQA: 一种基于掩码和选择的基于知识的视觉问答框架
机构 * Zhejiang University, Hangzhou, China(浙江大学, 杭州, 中国) ; Alibaba Group, Hangzhou, China(阿里巴巴集团, 杭州, 中国)
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 MaS-VQA通过结合显式知识过滤与隐式知识推理,提升基于知识的视觉问答任务的准确性和鲁棒性。
零样本系统用于体积CT和MRI图像的自动身体区域检测
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 本文提出零样本系统用于体积CT和MRI图像自动身体区域检测,通过预训练模型实现无监督分割,展示了基于规则和多模态语言模型的性能对比。
Comments 8 pages, 5 figures, 5 tables
无需训练的上下文取证链用于图像篡改检测与定位
专题命中 其他多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 ICFC提出一种无需训练的多模态大语言模型框架,用于图像篡改检测与定位,通过可解释的推理流程实现高效且准确的图像分析。
Comments This version was uploaded in error and contains misleading information found in an early draft. The manuscript requires extensive and long-term revisions
机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ; Beijing Normal University(北京师范大学)
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments 25 pages, 9 figures, 17 tables
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
机构 * School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) ; Beijing Digital Native Digital City Research Center(北京数字原生数字城市研究院) ; School of Computing, The University of Georgia(佐治亚大学计算机学院) ; School of Computer and Communication Engineering, University of Science and Technology Beijing(北京科技大学计算机与通信工程学院) ; Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学(深圳)计算机科学与工程系)
专题命中 其他多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI
Comments Accepted by TNNLS2025
机构 * KAIST, South Korea(韩国加尔文科学技术院) ; Auburn University, US(美国阿肯色大学)
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments Under Review
专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.AI
专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
Comments 35 pages, 18 figures
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments working in progress
专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL
Comments 9 pages
专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
Comments ACM MM2021
专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.CL
Comments Accepted by ECCV 2020. Code is available at https://github.com/spyflying/LSCM-Refseg
你所需要的只是对正常情况建模:多模态网络物理系统中异常检测的联合潜在聚类
机构 * Holon Institute of Technology (HIT)(霍隆技术学院) ; Afeka Academic College of Engineering(阿法卡工程学院)
专题命中 其他多模态 :multimodal(title)
AI总结 研究多模态网络物理系统异常检测,提出联合潜在聚类方法,通过MIIM假设集、公平协议及潜在评分建模正常行为,在三个真实数据集上表现优异,优于其他深度检测器。
Comments 17 pages, 1 figure
一种用于仿生空中机器人的多模态倾转机翼框架
机构 * Imperial College London(伦敦帝国理工学院)
专题命中 其他多模态 :multimodal(title)
AI总结 提出一种可切换悬停、高速前飞和高效滑翔的多模态倾转机翼框架,通过双独立扑翼推力矢量控制增强机动性,并采用混合苏格兰轭扑动机构实现宽扑动角度以利用拍合效应。
Comments 18 pages, 23 figures