Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style
AI 是否能像艺术史家一样看?解析视觉语言模型如何识别艺术风格
Marvin Limpijankit, Milad Alshomary, Yassin Oulad Daoud, Amith Ananthram, Tim Trombley, Emily L. Spratt, Anna Filonenko, Hannah Pivo, Elias Stengel-Eskin, Mohit Bansal, Noam M. Elcott, Kathleen McKeown
机构
*
Columbia University, Department of Computer Science(哥伦比亚大学计算机科学系)
;
Columbia University, Department of Art History & Archaeology(哥伦比亚大学艺术史与考古系)
;
University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
UNC Chapel Hill(北卡罗来纳大学教堂山分校)
专题命中
视觉问答
:vision language model(title);visual question answering(abstract);分类 cs.CV、cs.AI
V2X-QA: A Comprehensive Reasoning Dataset and Benchmark for Multimodal Large Language Models in Autonomous Driving Across Ego, Infrastructure, and Cooperative Views
V2X-QA:面向自动驾驶多视角的多模态大语言模型综合数据集与基准测试
Junwei You, Pei Li, Zhuoyu Jiang, Weizhe Tang, Zilin Huang, Rui Gan, Jiaxi Liu, Yan Zhao, Sikai Chen, Bin Ran
机构
*
Department of Civil and Environmental Engineering, University of Wisconsin–Madison(威斯康星大学麦迪逊分校土木与环境工程系)
;
Department of Civil and Architectural Engineering and Construction Management, University of Wyoming(怀俄明大学土木与建筑工程及施工管理系)
;
School of Transportation, Southeast University(东南大学交通学院)
专题命中
视觉问答
:multimodal large language model(title,abstract);分类 cs.CV、cs.AI
TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering
TPCL:面向鲁棒视觉问答的任务渐进课程学习
Ahmed Akl, Abdelwahed Khamis, Zhe Wang, Ali Cheraghian, Sara Khalifa, Kewen Wang
机构
*
School of Information and Communication Technology(信息与通信技术学院)
;
Data61, CSIRO Australia(Data61,澳大利亚联邦科学与工业研究组织)
;
School of Information Systems(信息系统学院)
机构
*
1 Institute of Automation, Chinese Academy of Sciences
;
2 School of Artificial Intelligence, University of Chinese Academy of Sciences
;
3 Centre for Artificial Intelligence
;
Robotics, Hong Kong Institute of Science \& Innovation, CAS
;
4 Independent Researcher
Saliency Guided Longitudinal Medical Visual Question Answering
基于显著性的纵向医学视觉问答
Jialin Wu, Xiaofeng Liu
机构
*
Dept. of Computer Science and Engineering University of California, San Diego(计算机科学与工程系,加州大学圣地亚哥分校)
;
Dept. of Radiology and Biomedical Imaging Yale University(放射学与生物医学成像系,耶鲁大学)
CMI-MTL: Cross-Mamba interaction based multi-task learning for medical visual question answering
Qiangguo Jin, Xianyao Zheng, Hui Cui, Changming Sun, Yuqi Fang, Cong Cong, Ran Su, Leyi Wei, Ping Xuan, Junbo Wang
机构
*
School of Software, Northwestern Polytechnical University, Shaanxi, China(西北工业大学软件学院)
;
Yangtze River Delta Research Institute of Northwestern Polytechnical University, Taicang, China(西北工业大学长江三角研究 institute)
;
Department of Computer Science and Information Technology, La Trobe University, Melbourne, Australia(拉筹伯大学计算机科学与信息技术系)
;
CSIRO Data61, Sydney, Australia(CSIRO Data61)
;
School of Intelligence Science and Technology, Nanjing University, Suzhou, China(南京大学智能科学与技术学院)
;
Australian Institute of Health Innovation (AIHI), Macquarie University, Australia(麦考瑞大学健康创新研究所)
;
School of Computer Software, College of Intelligence and Computing, Tianjin University, Tianjin, China(天津大学计算机软件学院)
;
Centre for Artificial Intelligence driven Drug Discovery, Faculty of Applied Science, Macao Polytechnic University, Macao Special Administrative Region of China(澳门理工学院人工智能驱动药物发现中心)
;
Department of Computer Science, School of Engineering, Shantou University, Guangdong, China(汕头大学计算机科学系)
DentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice
Zijie Meng, Jin Hao, Xiwei Dai, Yang Feng, Jiaxiang Liu, Bin Feng, Huikai Wu, Xiaotang Gai, Hengchuan Zhu, Tianxiang Hu, Yangyang Wu, Hongxia Xu, Jin Li, Jun Xiao, Xiaoqiang Liu, Joey Tianyi Zhou, Fudong Zhu, Zhihe Zhao, Lunguo Xia, Bing Fang, Jimeng Sun, Jian Wu, Zuozhu Liu
机构
*
Stomatology Hospital, School of Stomatology, Zhejiang University School of Medicine, Zhejiang University, Hangzhou(牙科医院,口腔医学院,浙江大学医学院,浙江大学,杭州)
;
College of Computer Science and Technology, Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University, Hangzhou(计算机科学与技术学院,浙江大学-伊利诺伊大学 Urbana-Champaign 院,浙江大学,杭州)
;
Department of Orthodontics, Shanghai Ninth People’s Hospital, College of Stomatology, Shanghai Jiao Tong University School of Medicine(正畸科,上海第九人民医院,口腔医学院,上海交通大学医学院)
;
Angelalign Technology Inc.(Angelalign 技术公司)
A Knowledge Noise Mitigation Framework for Knowledge-based Visual Question Answering
Zhiyue Liu, Sihang Liu, Jinyuan Liu, Xinru Zhang
机构
*
School of Computer, Electronics and Information(计算机、电子与信息学院)
;
Guangxi University(广西大学)
;
Guangxi Key Laboratory of Multimedia Communications and Network Technology(广西多媒体通信与网络技术重点实验室)