Red-teaming the Multimodal Reasoning: Jailbreaking Vision-Language Models via Cross-modal Entanglement Attacks
多模态推理的红队测试:通过跨模态纠缠攻击劫持视觉-语言模型
Yu Yan, Sheng Sun, Shengjia Cheng, Teli Liu, Mingfeng Li, Min Liu
机构
*
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
People’s Public Security University of China(中国人民公安大学)
机构
*
Shenzhen Key Lab for Advanced Motion Control and Modern Automation Equipments(深圳先进运动控制系统与现代自动化装备重点实验室)
;
Guangdong Provincial Key Laboratory of Intelligent Morphing Mechanisms and Adaptive Robotics(广东省智能变形机制与适应机器人省重点实验室)
;
School of Intelligence Science and Engineering(智能科学与工程学院)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
National Key Laboratory of Smart Farm Technologies and Systems(国家智能农业技术与系统重点实验室)
;
Autonomous Driving Center(自动驾驶中心)
;
Shanghai Utopilot Technology Co.Ltd.(上海驭目科技有限公司)
Grad2Reward: From Sparse Judgment to Dense Rewards for Improving Open-Ended LLM Reasoning
Grad2Reward: 从稀疏判断到密集奖励以提升开放性大语言模型推理
Zheng Zhang, Ao Lu, Yuanhao Zeng, Ziwei Shan, Jinjin Guo, Lufei Li, Yexin Li, Kan Ren
机构
*
School of Information Science and Technology, ShanghaiTech University(信息科学与技术学院,上海科技大学)
;
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)
ALIGN: Aligned Delegation with Performance Guarantees for Multi-Agent LLM Reasoning
ALIGN: 多智能体LLM推理中的对齐委托与性能保证
Tong Zhu, Baiting Chen, Jin Zhou, Hua Zhou, Sriram Sankararaman, Xiaowu Dai
机构
*
Department of Biostatistics, UCLA(生物统计学系,加州大学洛杉矶分校)
;
Department of Statistics and Data Science, UCLA(统计学与数据科学系,加州大学洛杉矶分校)
;
Department of Computer Science, UCLA(计算机科学系,加州大学洛杉矶分校)
;
Departments of Statistics and Data Science, and of Biostatistics, UCLA(统计学与数据科学系和生物统计学系,加州大学洛杉矶分校)
Jingcheng Deng, Liang Pang, Zihao Wei, Shicheng Xu, Zenghao Duan, Kun Xu, Yang Song, Huawei Shen, Xueqi Cheng
机构
*
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China(人工智能安全国家重点实验室,计算技术研究所,中国科学院,北京)
;
University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京)
LogicReward: Incentivizing LLM Reasoning via Step-Wise Logical Supervision
LogicReward: 通过逐步逻辑监督激励大语言模型推理
Jundong Xu, Hao Fei, Huichi Zhou, Xin Quan, Qijun Huang, Shengqiong Wu, William Yang Wang, Mong-Li Lee, Wynne Hsu
机构
*
National University of Singapore(新加坡国立大学)
;
University College London(伦敦大学学院)
;
University of Manchester(曼彻斯特大学)
;
University of Melbourne(墨尔本大学)
;
University of California, Santa Barbara(加州大学圣巴巴拉分校)
Physical Prompt Injection Attacks on Large Vision-Language Models
针对大视觉-语言模型的物理提示注入攻击
Chen Ling, Kai Hu, Hangcheng Liu, Xingshuo Han, Tianwei Zhang, Changhai Ou
机构
*
School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
;
College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院)