Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
Yuyang Hong, Jiaqi Gu, Qi Yang, Lubin Fan, Yue Wu, Ying Wang, Kun Ding, Shiming Xiang, Jieping Ye
机构
*
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS)
;
Alibaba Cloud Computing(阿里巴巴云计算)
专题命中
视觉问答
:visual language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI
Neglected Risks: The Disturbing Reality of Children's Images in Datasets and the Urgent Call for Accountability
Carlos Caetano, Gabriel O. dos Santos, Caio Petrucci, Artur Barros, Camila Laranjeira, Leo S. F. Ribeiro, Júlia F. de Mendonça, Jefersson A. dos Santos, Sandra Avila
机构
*
School of Computer Science, University of Sheffield(谢菲尔德大学计算机科学学院)
RegionMed-CLIP: A Region-Aware Multimodal Contrastive Learning Pre-trained Model for Medical Image Understanding
Tianchen Fang, Guiru Liu
专题命中
视觉问答
:vision language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI
CommentsUpon further review, we identified that our dataset requires optimization to ensure research reliability and accuracy. Additionally, considering the target journal's latest submission policies, we believe comprehensive manuscript revisions are necessary
Visual Enumeration Remains Challenging for Multimodal Generative AI
Alberto Testolin, Kuinan Hou, Marco Zorzi
机构
*
Department of General Psychology and Department of Mathematics University of Padova(帕多瓦大学心理学系和数学系)
;
Department of General Psychology University of Padova(帕多瓦大学心理学系)
;
Department of General Psychology and Padova Neuroscience Center University of Padova(帕多瓦大学心理学系和帕多瓦神经科学中心)
;
IRCSS San Camillo Hospital, Venice-Lido(威尼斯利多医院IRCSS桑卡莫医院)
机构
*
Department of Computer Science and Engineering, The Chinese University of Hong Kong(中国香港中文大学计算机科学与工程系)
;
School of Electronic Science and Engineering, Nanjing University(南京大学电子科学与工程学院)
;
School of Integrated Circuits, Peking University(北京大学集成电路学院)
;
School of Intergrated Circuits, Southeast University(东南大学集成电路学院)
;
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Department of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术系)
;
National Center of Technology Innovation for EDA(EDA技术创新国家中心)
专题命中
视觉问答
:multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG
Comments10 pages, 1 figure, 5 tables. To appear in ICCAD 2025
TextSquare: Scaling up Text-Centric Visual Instruction Tuning
Jingqun Tang, Chunhui Lin, Zhen Zhao, Shu Wei, Binghong Wu, Qi Liu, Yangfan He, Kuan Lu, Hao Feng, Yang Li, Siqi Wang, Lei Liao, Wei Shi, Yuliang Liu, Hao Liu, Yuan Xie, Xiang Bai, Can Huang
机构
*
ByteDance Inc.(字节跳动公司)
;
East China Normal University(东华大学)
;
Huazhong University of Science and Technology(华中科技大学)
;
University of Minnesota(明尼苏达大学)
;
Cornell University(康奈尔大学)
专题命中
视觉问答
:visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV、cs.LG
专题命中
视觉问答
:visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV、cs.LG
CommentsClarification note for the CVPR 2025 paper (FarSight). Prepared by a subset of the original authors; remaining co-authors are acknowledged in the text