Physical Prompt Injection Attacks on Large Vision-Language Models
针对大视觉-语言模型的物理提示注入攻击
Chen Ling, Kai Hu, Hangcheng Liu, Xingshuo Han, Tianwei Zhang, Changhai Ou
机构
*
School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
;
College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院)
Ask Me Again Differently: GRAS for Measuring Bias in Vision Language Models on Gender, Race, Age, and Skin Tone
再次提问:GRAS用于测量视觉语言模型在性别、种族、年龄和肤色上的偏见
Shaivi Malik, Hasnat Md Abdullah, Sriparna Saha, Amit Sheth
机构
*
Guru Gobind Singh Indraprastha University(古鲁·戈宾德·辛格印度教普拉斯塔大学)
;
AI Institute, University of South Carolina(南卡罗来纳大学人工智能研究所)
;
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Indian Institute of Technology Patna, India(印度理工学院帕纳布分校)
专题命中
视觉问答
:vision language model(title,abstract);visual question answering(abstract);分类 cs.CV
GRASP: Guided Region-Aware Sparse Prompting for Adapting MLLMs to Remote Sensing
GRASP: 基于区域感知的稀疏提示引导方法用于适应遥感图像的多模态大语言模型
Qigan Sun, Chaoning Zhang, Jianwei Zhang, Xudong Wang, Jiehui Xie, Pengcheng Zheng, Haoyu Wang, Sungyoung Lee, Chi-lok Andy Tai, Yang Yang, Heng Tao Shen
机构
*
School of Computing, Kyung Hee University(京畿大学计算机学院)
;
School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)
;
School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)
;
College of Computer Science and Information Engineering, Harbin Normal University(哈尔滨师范大学计算机科学与信息工程学院)
;
College of Professional and Continuing Education, The Hong Kong Polytechnic University(香港理工大学专业及继续教育学院)
;
School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)
专题命中
视觉问答
:visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV
机构
*
University of Adelaide(阿德莱德大学)
;
University of Liverpool(利物浦大学)
;
University of Technology Sydney(悉尼技术大学)
;
The Australian National University(澳大利亚国立大学)
;
La Trobe University(拉特罗布大学)
CommentsThe paper is withdrawn by the authors after discovering a flaw in the theoretical derivation presented in the Method section. This incorrect step leads to conclusions that are not supported by the corrected derivation. The authors plan to reconstruct the argument and will release an updated version once the issue is fully resolved
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding
X-LeBench:一个用于极长第一人称视频理解的基准数据集
Wenqi Zhou, Kai Cao, Hao Zheng, Yunze Liu, Xinyi Zheng, Miao Liu, Per Ola Kristensson, Walterio Mayol-Cuevas, Fan Zhang, Weizhe Lin, Junxiao Shen
机构
*
University of Bristol(布里斯托大学)
;
University of Manchester(曼彻斯特大学)
;
University of Cambridge(剑桥大学)
;
College of AI, Tsinghua University(清华大学人工智能学院)
;
Meta
;
Memories.ai Research(Memories.ai研究)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning
超越结果验证:用于结构化推理的可验证过程奖励模型
Massimiliano Pronesti, Anya Belz, Yufang Hou
机构
*
IBM Research Europe - Ireland(IBM欧洲研究院-爱尔兰)
;
Dublin City University(都柏林城市大学)
;
IT:U Interdisciplinary Transformation University Austria(IT:U跨学科转型大学奥地利)
CommentsInitial Version, Pending Updates. We welcome any feedback and suggestions for improvement. Please feel free to contact us at an.hongjun@foxmail.com