机构
*
Peking University(北京大学)
;
Mohamed bin Zayed University of Artificial Intelligence(莫扎伊德大学人工智能学院)
;
National University of Singapore(新加坡国立大学)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
University of Science and Technology of China(中国科学技术大学)
;
Cornell University(康奈尔大学)
;
Hong Kong Polytechnic University(香港理工大学)
;
City University of Hong Kong(香港城市大学)
专题命中
视觉推理
:visual reasoning(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams
Event-VStream: 基于事件驱动的长视频流实时理解
Zhenghui Guo, Yuanbin Man, Junyuan Sheng, Bowen Lin, Ahmed Ahmed, Bo Jiang, Boyuan Zhang, Miao Yin, Sian Jin, Omprakash Gnawal, Chengming Zhang
机构
*
University of Houston(德克萨斯大学休斯顿分校)
;
The University of Texas at Arlington(德克萨斯理工大学)
;
Indiana University Bloomington(印第安纳大学布卢明顿分校)
;
Temple University(特拉华大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV、cs.AI
The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection
范式转变:大型视觉语言模型在多模态虚假新闻检测中的全面调查
Wei Ai, Yilong Tan, Yuntao Shou, Tao Meng, Haowen Chen, Zhixiong He, Keqin Li
机构
*
College of Computer and Mathematics, Central South University of Forestry and Technology(计算机与数学学院,中央南大学林业科技学院)
;
College of Computer Science and Electronic Engineering, Hunan University(计算机科学与电子工程学院,湖南大学)
;
College of Economics and Management, Central South University of Forestry and Technology(经济管理学院,中央南大学林业科技学院)
;
Department of Computer Science, State University of New York(计算机科学系,纽约州立大学)
专题命中
视觉定位与Grounding
:vision language model(title);vision-language model(abstract);分类 cs.CV、cs.AI
Grounding Large Language Models in Reaction Knowledge Graphs for Synthesis Retrieval
将反应知识图谱接地于大语言模型以实现合成检索
Olga Bunkova, Lorenzo Di Fruscia, Sophia Rupprecht, Artur M. Schweidtmann, Marcel J. T. Reinders, Jana M. Weber
机构
*
Department of Intelligent Systems, Delft University of Technology(智能系统系,代尔夫特理工大学)
;
Department of Chemical Engineering, Delft University of Technology(化学工程系,代尔夫特理工大学)
DevPrompt: Deviation-Based Prompt Learning for One-Normal ShotImage Anomaly Detection
DevPrompt: 基于偏差的提示学习用于单正常样本图像异常检测
Morteza Poudineh, Marc Lalonde
机构
*
Department of Electrical and Computer Engineering, Concordia University(电气与计算机工程系,康科迪亚大学)
;
Computer Research Institute of Montreal (CRIM)(蒙特利尔计算机研究 institute)
Crafting Adversarial Inputs for Large Vision-Language Models Using Black-Box Optimization
为大型视觉-语言模型设计对抗输入使用黑盒优化
Jiwei Guan, Haibo Jin, Haohan Wang
机构
*
School of Computing, Macquarie University(麦考瑞大学计算机学院)
;
School of Information Sciences, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校信息科学学院)
Beyond Visual Safety: Jailbreaking Multimodal Large Language Models for Harmful Image Generation via Semantic-Agnostic Inputs
超越视觉安全:通过语义无关输入对多模态大语言模型进行有害图像生成的劫持
Mingyu Yu, Lana Liu, Zhehao Zhao, Wei Wang, Sujuan Qin
机构
*
State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(网络与交换技术国家重点实验室,北京邮电大学)
;
School of Cyberspace Security, Beijing University of Posts and Telecommunications(网络安全学院,北京邮电大学)
专题命中
幻觉与鲁棒性
:multimodal large language model(title,abstract);分类 cs.CV、cs.AI
Hallucination Mitigating for Medical Report Generation
缓解医疗报告生成中的幻觉
Ruoqing Zhao, Runze Xia, Piji Li
机构
*
College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院)
;
MIIT Key Laboratory of Pattern Analysis and Machine Intelligence(信息产业部模式分析与机器智能重点实验室)
;
The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(教育部脑机智能技术重点实验室)
Comments[fixed typos in equations] accepted to ICCV2023 Poster; some figures are not supported when viewed online, please download the file and view locally