ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning
机构 * Centre for Smart Health, School of Nursing, The Hong Kong Polytechnic University(智能健康研究中心、护理学院、香港理工大学) ; Department of Language Science and Technology, The Hong Kong Polytechnic University(语言科学与技术系、香港理工大学) ; Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(电子与电气工程系、香港理工大学)
专题命中 推理评测 :reasoning(title,abstract);chain-of-thought(title,abstract);CoT(abstract);分类 cs.AI
Comments 9 pages, 2 figures