Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks
出血路径:LLM隐藏状态中的判别能力消失加剧了 jailbreak 攻击
Yingjie Zhang, Tong Liu, Zhe Zhao, Guozhu Meng, Kai Chen
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
Ant Group(蚂蚁集团)
Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection
通过双层图异常检测实现LLM多智能体系统的可解释性与细粒度防护
Junjun Pan, Yixin Liu, Rui Miao, Kaize Ding, Yu Zheng, Quoc Viet Hung Nguyen, Alan Wee-Chung Liew, Shirui Pan
机构
*
School of Information and Communication Technology, Griffith University, Australia(格里菲斯大学信息与通信技术学院)
;
School of Artificial Intelligence, Jilin University, China(吉林大学人工智能学院)
;
Department of Statistics and Data Science, Northwestern University, USA(西北大学统计与数据科学系)