SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
SHE:面向大语言模型智能体的轨迹驱动安全管控机制演化
Wanying Qu, Qinghua Mao, Yu Li, Jiyao Liu, Xin Zhang, Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Hu, Dongrui Liu
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Fudan University(复旦大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
机构
*
University of Colorado Boulder(科罗拉多大学博尔德分校)
;
University of Central Florida(中佛罗里达大学)
;
University of Maryland College Park(马里兰大学帕克分校)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
机构
*
School of Software Technology, Zhejiang University(浙江大学软件学院)
;
State Key Laboratory of CAD&CG, Zhejiang University(浙江大学CAD&CG国家重点实验室)
;
Stomatology Hospital, Zhejiang University School of Medicine(浙江大学医学院附属口腔医院)
TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent
TAF-MED:声明自我治疗意图下大语言模型的多轮安全拒绝崩溃
Waleed Jamil, Raphael Schmitt
机构
*
Independent Researcher(独立研究者)
;
School of Computation, Information and Technology, Technical University of Munich(慕尼黑工业大学计算、信息与技术学院)
;
Institute of General Practice, Faculty of Medicine and Medical Center, University of Freiburg(弗莱堡大学医学中心医学院普通实践研究所)