Toward Autonomous Laboratory Safety Monitoring with Vision Language Models: Learning to See Hazards Through Scene Structure
迈向自主实验室安全监控:通过场景结构学习识别危险
Trishna Chakraborty, Udita Ghosh, Aldair Ernesto Gongora, Ruben Glatt, Yue Dong, Jiachen Li, Amit K. Roy-Chowdhury, Chengyu Song
机构
*
Electrical and Computer Engineering, University of California, Riverside, USA(加州大学河滨分校电气与计算机工程系)
;
Computer Science and Engineering, University of California, Riverside, USA(加州大学河滨分校计算机科学与工程系)
;
Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)
专题命中
幻觉与鲁棒性
:vision language model(title,abstract);VLM(abstract);分类 cs.CV、cs.LG
机构
*
School of Computer Science(计算机科学学院)
;
Artificial Intelligence, Wuhan University of Technology, Hubei 430070, China(人工智能,武汉科技大学,湖北430070,中国)
;
School of Computer Science, Wuhan University, Hubei 430072, China(计算机科学学院,武汉大学,湖北430072,中国)
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
迈向通用且可迁移的视觉-语言模型劫持攻击
Kaiyuan Cui, Yige Li, Yutao Wu, Xingjun Ma, Sarah Erfani, Christopher Leckie, Hanxun Huang
机构
*
School of Computing and Information Systems, The University of Melbourne, Australia(墨尔本大学计算机与信息系)
;
School of Computing and Information Systems, Singapore Management University, Singapore(新加坡管理大学计算机与信息系)
;
School of Information Technology, Deakin University, Australia(德肯大学信息科技系)
;
Institute of Trustworthy Embodied AI, Fudan University, China(复旦大学可信具身人工智能研究所)
机构
*
School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院)
;
Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(复旦大学计算机学院智能信息处理重点实验室)
;
Department of Mathematics and Applications, University of Naples Federico II(那不勒斯费德里克二世大学应用数学系)
;
State Key Laboratory of Advanced Rail Autonomous Operation, Beijing Jiaotong University(北京交通大学先进轨道交通自主运行国家重点实验室)
ClueTracer: Question-to-Vision Clue Tracing for Training-Free Hallucination Suppression in Multimodal Reasoning
ClueTracer: 问题到视觉线索追踪用于无训练 hallucination 抑制在多模态推理
Gongli Xi, Kun Wang, Zeming Gao, Huahui Yi, Haolang Lu, Ye Tian, Wendong Wang
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Nanyang Technological University(南洋理工大学)
;
West China Biomedical Big Data Center(西京生物大数据中心)
Semantically Guided Dynamic Visual Prototype Refinement for Compositional Zero-Shot Learning
语义引导的动态视觉原型精炼用于组合零样本学习
Zhong Peng, Yishi Xu, Gerong Wang, Wenchao Chen, Bo Chen, Jing Zhang, Hongwei Liu
机构
*
National Key Laboratory of Radar Signal Processing, Xidian University(雷达信号处理国家级重点实验室,西安电子科技大学)
;
Research Institute of Systems Engineering, Academy of Military Science(系统工程研究所,军事科学院)
HalluHard: A Hard Multi-Turn Hallucination Benchmark
HalluHard: 一种具有挑战性的多轮 hallucination 评估基准
Dongyang Fan, Sebastien Delsad, Nicolas Flammarion, Maksym Andriushchenko
机构
*
EPFL(苏黎世联邦理工学院)
;
ELLIS Institute Tübingen(图宾根ELLIS研究所)
;
Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
;
Tübingen AI Center(图宾根人工智能中心)
机构
*
School of Electronic Engineering, Xidian University, Xi'an, China(西安电子科技大学电子工程学院)
;
School of Computer Science and Technology, Xidian University, Xi'an, China(西安电子科技大学计算机科学与技术学院)
;
College of Computer and Information, Hohai University, Nanjing, China(河海大学计算机与信息学院)
;
Institute for Infocomm Research, A*STAR, Singapore(新加坡资讯研究院)
Hallucination-Resistant Relation Extraction via Dependency-Aware Sentence Simplification and Two-tiered Hierarchical Refinement
通过依赖感知句子简化和双层级层次精炼实现抗幻觉的关系抽取
Yupei Yang, Fan Feng, Lin Yang, Wanxi Deng, Lin Qu, Biwei Huang, Shikui Tu, Lei Xu
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
University of California San Diego(加州大学圣地亚哥分校)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Alibaba Group(阿里巴巴集团)
MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
MLLMEraser: 通过激活引导在多模态大语言模型中实现测试时遗忘
Chenlu Ding, Jiancan Wu, Leheng Sheng, Fan Zhang, Yancheng Yuan, Xiang Wang, Xiangnan He
机构
*
University of Science and Technology of China(中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Huawei(华为)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);LLaVA(abstract);MLLM(abstract);分类 cs.AI、cs.LG
SGHA-Attack: Semantic-Guided Hierarchical Alignment for Transferable Targeted Attacks on Vision-Language Models
SGHA-Attack:语义引导的层次对齐用于视觉-语言模型的可迁移定向攻击
Haobo Wang, Weiqi Luo, Xiaojun Jia, Xiaochun Cao
机构
*
Guangdong Province Key Lab of Information Security Technology, and School of Computer Science and Engineering, Sun Yat-sen University(广东信息安全技术重点实验室,计算机科学与工程学院,中山大学)
;
Nanyang Technological University(南洋理工大学)
;
School of Cyber Science and Technology, Shenzhen Campus, Sun Yat-sen University(网络安全科学与技术学院,深圳校区,中山大学)
A Survey of Token Compression for Efficient Multimodal Large Language Models
多模态大语言模型高效性中的标记压缩综述
Kele Shao, Keda Tao, Kejia Zhang, Sicheng Feng, Mu Cai, Yuzhang Shang, Haoxuan You, Can Qin, Yang Sui, Huan Wang
机构
*
Zhejiang University(浙江大学)
;
Westlake University(西湖大学)
;
Xiamen University(厦门大学)
;
National University of Singapore(新加坡国立大学)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
;
University of Central Florida(佛罗里达大学)
;
Salesforce AI Research(Salesforce AI研究)
;
Rice University(德克萨斯大学)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
Model Specific Task Similarity for Vision Language Model Selection via Layer Conductance
为视觉语言模型选择而基于模型特定任务相似性的层导电性
Wei Yang, Hong Xie, Tao Tan, Xin Li, Defu Lian, Enhong Chen
机构
*
School of Computer Science and Technology, University of Science and Technology of China(计算机科学与技术学院,科学技术大学)
;
School of AI and Data Science, University of Science and Technology of China(人工智能与数据科学学院,科学技术大学)
专题命中
VLM训练与架构
:vision language model(title);vision-language model(abstract);分类 cs.AI