CommentsCopyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
Visual Enumeration Remains Challenging for Multimodal Generative AI
Alberto Testolin, Kuinan Hou, Marco Zorzi
机构
*
Department of General Psychology and Department of Mathematics University of Padova(帕多瓦大学心理学系和数学系)
;
Department of General Psychology University of Padova(帕多瓦大学心理学系)
;
Department of General Psychology and Padova Neuroscience Center University of Padova(帕多瓦大学心理学系和帕多瓦神经科学中心)
;
IRCSS San Camillo Hospital, Venice-Lido(威尼斯利多医院IRCSS桑卡莫医院)
机构
*
Department of Computer Science and Engineering, The Chinese University of Hong Kong(中国香港中文大学计算机科学与工程系)
;
School of Electronic Science and Engineering, Nanjing University(南京大学电子科学与工程学院)
;
School of Integrated Circuits, Peking University(北京大学集成电路学院)
;
School of Intergrated Circuits, Southeast University(东南大学集成电路学院)
;
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Department of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术系)
;
National Center of Technology Innovation for EDA(EDA技术创新国家中心)
专题命中
视觉问答
:multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG
Comments10 pages, 1 figure, 5 tables. To appear in ICCAD 2025
机构
*
University of Science and Technology of China(科学技术大学)
;
University of Adelaide(阿德莱德大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
专题命中
视觉推理
:visual language model(title);vision-language model(abstract);VLM(abstract);分类 cs.AI
Comments8 pages, 3 figures, this paper has been accepted by ACM MM 2025
机构
*
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
Wuhan National Laboratory for Optoelectronics, Huazhong University of Science and Technology(华中科技大学光电研究院)
;
Meta Reality Lab(Meta现实实验室)
;
Xi’an Jiao Tong University(西安交通大学)
;
National University of Singapore(新加坡国立大学)
;
The Department of Systems Hub, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学系统枢纽部门(广州))
FreeQ-Graph: Free-form Querying with Semantic Consistent Scene Graph for 3D Scene Understanding
Chenlu Zhan, Yufei Zhang, Gaoang Wang, Hongwei Wang
机构
*
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
College of Biomedical Engineering and Instrument Science, Zhejiang University(浙江大学生物医学工程与仪器科学学院)
;
Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University(浙江大学-伊利诺伊大学厄巴纳-香槟分校联合学院)
Zero-shot Performance of Generative AI in Brazilian Portuguese Medical Exam
Cesar Augusto Madid Truyts, Amanda Gomes Rabelo, Gabriel Mesquita de Souza, Daniel Scaldaferri Lages, Adriano Jose Pereira, Uri Adrian Prync Flato, Eduardo Pontes dos Reis, Joaquim Edson Vieira, Paulo Sergio Panse Silveira, Edson Amaro Junior
机构
*
Einstein Global Advanced Technologies for Equity(埃因斯坦全球先进科技以公平为宗旨)
;
Hospital Israelita Albert Einstein(埃因斯坦医院)
;
Departamento de Pacientes Graves(重症患者部门)
;
Stanford Center for Artificial Intelligence in Medicine and Imaging(斯坦福大学医学与成像人工智能中心)
;
Departmento de Cirurgia(外科部门)
;
Faculdade de Medicina, Universidade de São Paulo(圣保罗大学医学院)
;
Faculdade Israelita de Ciências da Saúde Albert Einstein(埃因斯坦以色列健康科学学院)
专题命中
视觉推理
:multimodal large language model(abstract)
ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition
Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He
机构
*
South China University of Technology(华南理工大学)
;
Guangdong Engineering Center for Large Model and GenAI Technology(广东省大模型与生成式人工智能技术工程中心)
;
State Key Laboratory of Subtropical Building and Urban Science(亚热带建筑科学国家重点实验室)
;
Ministry of Education Key Laboratory of Big Data and Intelligent Robot(教育部大数据与智能机器人重点实验室)
;
Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information(广东省计算智能与网络信息重点实验室)
;
Singapore Management University(新加坡国立大学)
CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning
George Ibrahim, Rita Ramos, Yova Kementchedjhieva
机构
*
Department of Natural Language Processing, MBZUAI(自然语言处理部门,MBZUAI)
;
INESC-ID, Instituto Superior Técnico, University of Lisbon(INESC-ID,理工学院,里斯本大学)
EventVAD: Training-Free Event-Aware Video Anomaly Detection
Yihua Shao, Haojin He, Sijie Li, Siyu Chen, Xinwei Long, Fanhu Zeng, Yuxuan Fan, Muyang Zhang, Ziyang Yan, Ao Ma, Xiaochen Wang, Hao Tang, Yan Wang, Shuyan Li
机构
*
Peking University(北京大学)
;
Guangdong University of Technology(广东工业大学)
;
The University of Sheffield(谢菲尔德大学)
;
University of Science and Technology Beijing(北京科技大学)
;
Tsinghua University(清华大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
Nanjing University(南京大学)
;
University of Trento(特伦特大学)
;
Queen's University Belfast(贝尔法斯特女王大学)