AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
AutoViVQA:一个大规模自动构建的数据集用于越南语视觉问答
Nguyen Anh Tuong, Phan Ba Duc, Nguyen Trung Quoc, Tran Dac Thinh, Dang Duy Lan, Nguyen Quoc Thinh, Tung Le
机构
*
Faculty of Information Technology, University of Science, VNU-HCM(越南国家大学胡志明市分校信息科技学院)
;
Vietnam National University, Ho Chi Minh City(越南国家大学胡志明市分校)
DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAs
DynamicGTR: 利用图拓扑表示偏好提升视觉语言模型在图问答中的能力
Yanbin Wei, Jiangyue Yan, Chun Kang, Yang Chen, Hua Liu, James Kwok, Yu Zhang
机构
*
Southern University of Science and Technology(南方科技大学)
;
Hong Kong University of Science and Technology(香港理工大学)
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Beihang University(北京航空航天大学)
FaithSCAN: Model-Driven Single-Pass Hallucination Detection for Faithful Visual Question Answering
FaithSCAN: 基于模型的单次幻觉检测用于可靠的视觉问答
Chaodong Tong, Qi Zhang, Chen Li, Lei Jiang, Yanbing Liu
机构
*
Institute of Information Engineering, Chinese Academy of Sciences (CAS) and the School of Cyber Security, University of CAS(信息工程研究所、中国科学院(CAS)和安全学院、CAS大学)
Physical Prompt Injection Attacks on Large Vision-Language Models
针对大视觉-语言模型的物理提示注入攻击
Chen Ling, Kai Hu, Hangcheng Liu, Xingshuo Han, Tianwei Zhang, Changhai Ou
机构
*
School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
;
College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院)
SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
SurgMLLMBench: 一个用于手术场景理解的多模态大语言模型基准数据集
Tae-Min Choi, Tae Kyeong Jeong, Garam Kim, Jaemin Lee, Yeongyoon Koh, In Cheul Choi, Jae-Ho Chung, Jong Woong Park, Juyoun Park
机构
*
Samsung Research(三星研究所)
;
Center for Humanoid Research, Korea Institute of Science and Technology(人类研究学院,韩国科学技术院)
;
Department of plastic surgery, College of medicine, Korea University(医学院整形外科部,韩国大学)
;
Department of orthopedic surgery, College of medicine, Korea University(医学院骨科部,韩国大学)
专题命中
视觉问答
:multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI
机构
*
College of Computer, National University of Defense Technology(国防科技大学计算机学院)
;
School of Computer and Information Engineering, Hefei University of Technology(合肥工业大学计算机与信息工程学院)
;
CCNU(中国地质大学)
专题命中
视觉问答
:multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV、cs.AI
CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
Hongyong Han, Wei Wang, Gaowei Zhang, Mingjie Li, Yi Wang
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Technology Innovation Center for South China Sea Remote Sensing, Surveying and Mapping Collaborative Application, Ministry of Natural Resources(自然资源部南海遥感测绘协同应用技术创新中心)
;
South China Sea Development Research Institute, Ministry of Natural Resources(自然资源部南海发展研究 institute)
;
Inspur Computer Technology Co., Ltd(Inspur 计算机技术有限公司)
;
Shandong Key Laboratory of Advanced Computing(山东先进计算重点实验室)
SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
Yan Zhang, Jiaqing Lin, Miao Zhang, Kui Xiao, Xiaoju Hou, Yue Zhao, Zhifei Li
机构
*
School of Computer Science, Hubei University, Wuhan, China(湖北大学计算机学院)
;
Hubei Key Laboratory of Big Data Intelligent Analysis and Application (Hubei University)(湖北大数据智能分析与应用重点实验室)
;
Key Laboratory of Intelligent Sensing System and Security (Hubei University)(智能传感系统与安全重点实验室)
;
Institute of Vocational Education, Guangdong Industry Polytechnic University(广东行业职业大学职业教育学院)
;
Shandong Police College(山东警察学院)
专题命中
视觉问答
:visual question answering(title,abstract);visual language model(abstract);分类 cs.CV、cs.AI
CommentsACCEPTED as a FULL PAPER for the Research Track at International Conference on Database Systems for Advanced Applications 2025
机构
*
AI VIETNAM Lab(AI越南实验室)
;
Carnegie Mellon University(卡内基梅隆大学)
;
University of Wisconsin - Madison(威斯康星大学麦迪逊分校)
;
University of Pittsburgh(匹兹堡大学)
;
University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)
;
Northwestern University(西北大学)
机构
*
Massachusetts Institute of Technology(麻省理工学院)
;
Amazon Web Services(亚马逊网络服务)
专题命中
视觉问答
:VLM(title,abstract);vision language model(abstract);分类 cs.CV、cs.AI
Comments8 pages, 5 figures, accepted to the 11th IEEE International Workshop on Computer Vision in Sports (CVSports) at CVPR 2025; supplementary appendix included
Sample then Identify: A General Framework for Risk Control and Assessment in Multimodal Large Language Models
Qingni Wang, Tiantian Geng, Zhiyuan Wang, Teng Wang, Bo Fu, Feng Zheng
机构
*
University of Electronic Science and Technology of China(电子科技大学)
;
Southern University of Science and Technology(南方科技大学)
;
University of Birmingham(伯明翰大学)
;
The University of Hong Kong(香港大学)
专题命中
视觉问答
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI、cs.LG
Beyond the Hype: A dispassionate look at vision-language models in medical scenario
Yang Nan, Huichi Zhou, Xiaodan Xing, Guang Yang
机构
*
Bioengineering Department and Imperial-X, Imperial College London(生物工程部门和Imperial-X,帝国理工学院伦敦分校)
;
GSK, Artificial Intelligence and Machine Learning(GSK,人工智能与机器学习)
;
Cardiovascular Research Centre, Royal Brompton Hospital(心血管研究中心,皇家布里托尼医院)
;
School of Biomedical Engineering and Imaging Sciences, King’s College London(生物医学工程与成像科学学院,国王学院伦敦分校)