ORBIT: Training-Free Multi-Attribute Behavioral Steering via Orthogonal Subspace Rotation
ORBIT: 通过正交子空间旋转实现无训练的多属性行为引导
Narges Ghasemi, Amir Ziashahabi, Salman Avestimehr, Jonathan May
机构
*
Information Sciences Institute, University of Southern California(南加州大学信息科学研究所)
;
Department of Electrical and Computer Engineering, University of Southern California(南加州大学电气与计算机工程系)
When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs
当技能与安全相遇:对技能合并大语言模型的自适应越狱鲁棒性进行基准测试与表征
Yu Ma, Hongli Shi, Jing Li, Xinran Xu, Weiwei Hou
机构
*
Google(谷歌公司)
;
University of New South Wales(新南威尔士大学)
;
University of Technology Sydney(悉尼科技大学)
;
Zhejiang University(浙江大学)
;
Australian National University(澳大利亚国立大学)
Illusion of Alignment: Detecting Hidden Disagreement in Collaborative Dialogue
对齐的错觉:检测协同对话中的隐藏分歧
Kaiming Liu, Fuwen Luo, Ziyue Wang, Jinrui Ju, Yuxuan Liu, Xuanyu Lei, Yunghwei Lai, Peng Li, Yang Liu
机构
*
College of AI, Tsinghua University(清华大学人工智能学院)
;
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
;
Institute for AI, Tsinghua University(清华大学人工智能研究院)
ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
通道卫士:安全模型无法组合成安全的多智能体系统
Elias Hossain, Md Mehedi Hasan Nipu, Fatema Tuj Johora Faria, Tasfia Nuzhat Ornee, Maleeha Sheikh
机构
*
College of Engineering and Computer Science, University of Central Florida(工程与计算机科学学院,中央佛罗里达大学)
;
Department of Computer Science and Engineering, North South University(计算机科学与工程系,北南大学)
;
Computer Science and Engineering, Ahsanullah University of Science and Technology(计算机科学与工程,阿沙努拉科学与技术大学)
;
Department of Electrical and Computer Engineering, Purdue University Fort Wayne(电气与计算机工程系,普渡大学弗拉特沃恩分校)