机构
*
University of South Florida(佛罗里达南大学)
;
Missouri University of Science and Technology(密苏里科技大学)
;
University of Alabama(阿拉巴马大学)
;
Florida International University(佛罗里达国际大学)
;
University of Cincinnati(辛辛那提大学)
;
George Mason University(乔治·梅森大学)
Authority Inversion in LLM-Mediated Ubiquitous Systems: When Models Trust Users Over Sensors
LLM介导的普适系统中的权威倒置:当模型信任用户胜过传感器
Long Zhang, Zi-bo Qin, Wei-neng Chen
机构
*
School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院)
;
School of Computer Science(计算机科学学院)
;
Engineering, South China University of Technology(华南理工大学工程学院)
Lei Wang, Syuan-Hao Li, Yongsheng Gao, Piotr Koniusz
机构
*
School of Engineering and Built Environment, Electrical and Electronic Engineering, Griffith University(工程与建筑环境学院,电气与电子工程学院,格里菲斯大学)
;
School of Computer Science and Engineering, University of New South Wales(计算机科学与工程学院,新南威尔士大学)
Universal Boosts, Specific Suppressors: Sparse Autoencoder Steering of Medical Vision-Language Models
通用增强,特定抑制:基于稀疏自编码器引导的医学视觉语言模型
Farhad Nooralahzadeh, Benjamin Gundersen, Nicolas Deperrois, Hidetoshi Matsuom, Mizuho Nishio, Thomas Frauenfelder, Ahmed Allam, Christian Blüthgen, Michael Moor, Michael Krauthammer
机构
*
University of Zurich and University Hospital of Zurich(苏黎世大学及苏黎世大学医院)
;
Kobe University(Kobe大学)
;
ETH AI Center(苏黎世联邦理工学院人工智能中心)
;
ETH Zurich(苏黎世联邦理工学院)
;
Stanford University(斯坦福大学)
;
Zurich University of Applied Sciences(苏黎世应用科学大学)
机构
*
Hasso Plattner Institute, University of Potsdam(哈索普兰特纳研究所,波茨坦大学)
;
TU Darmstadt, Secure Mobile Networking Lab(德累斯顿技术大学,安全移动网络实验室)
;
IMDEA Networks Institute, Madrid, Spain(IMDEA网络研究所,马德里,西班牙)
机构
*
Faculty of Engineering, Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学工程学院、计算机科学与工程系)
;
Artificial Intelligence Innovation and Incubation Institute, Fudan University(复旦大学人工智能创新与孵化院)
;
Shanghai Academy of AI for Science(上海人工智能科学研究院)
AI总结
本文通过 Hateful Memes Challenge 数据集系统分析 GPT-4o mini 在多模态仇恨言论检测中的安全架构,发现并实验验证了“单模态瓶颈”缺陷,即上下文无关的安全过滤器会优先阻断多模态推理,导致误报。
CommentsThis paper reports preliminary findings from a small-scale study whose sample size is insufficient to support the stated conclusions. The authors are withdrawing it to conduct a more comprehensive evaluation
From Automation to Collaboration: Human-in-the-Loop Methods for Safe and Trustworthy NLP
从自动化到协作:面向安全可信NLP的人机协同方法
Most. Sharmin Sultana Samu, MD. Tanvir Ahmed Seum, Md. Rakibul Islam
机构
*
Department of Computer Science and Engineering, BRAC University(布拉克大学计算机科学与工程系)
;
Department of Electrical and Electronic Engineering, Rajshahi University of Engineering and Technology(拉贾沙希工程与技术大学电子与电气工程系)