HumorReject: Decoupling LLM Safety from Refusal Prefix via A Little Humor
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.LG
机构 * PKU Alignment Team, Peking University(北京大学对齐团队) ; LLM Safety Centre, Beijing Academy of Artificial Intelligence(北京人工智能研究院大语言模型安全中心)
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.AI
机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) ; King Abdullah University of Science and Technology(卡塔尔国王大学科学与技术研究院) ; University of Oxford(牛津大学)
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.AI
机构 * Macquarie University(麦觉里大学)
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.CL
Comments Preprint
机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.CL
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.AI
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.LG
专题命中 安全训练 :alignment(title,abstract);safety(abstract);分类 cs.AI
专题命中 安全训练 :safety(title,abstract);AI safety(abstract);分类 cs.CY
机构 * University of Strasbourg, CNRS, INSERM, ICube, UMR7357, Strasbourg, France(斯特拉斯堡大学,法国国家科学研究中心,法国国家卫生研究院,ICube,UMR7357,斯特拉斯堡,法国) ; Fondazione Policlinico Universitario A. Gemelli IRCCS, Università Cattolica del Sacro Cuore, Rome, Italy(A. Gemelli IRCCS大学医院,罗马,意大利) ; CAMP, Technische Universität München, Munich, Germany(慕尼黑技术大学,德国) ; Institute of Image-Guided Surgery, IHU Strasbourg, Strasbourg, France(影像引导手术研究所,斯特拉斯堡,法国)
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.AI
机构 * Georgia Tech(佐治亚理工学院) ; Mila ; UC Berkeley(加州大学伯克利分校) ; Stanford(斯坦福大学) ; MBZUAI(穆桑大学人工智能研究所) ; McGill(麦吉尔大学)
专题命中 安全训练 :safety(title,abstract);AI safety(abstract);分类 cs.AI
机构 * University of Science and Technology of China(中国科学技术大学)
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.AI
Comments 16 pages, 10 figures
专题命中 安全训练 :alignment(title,abstract);safety(abstract);分类 cs.AI
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.CY
Comments 75 pages, 41 tables/figures
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.AI
Comments Extended version of paper accepted to AAAI 2025. 14 pages, 6 figures
专题命中 安全训练 :safety(title,abstract);jailbreak(abstract);分类 cs.CL
Comments NeurIPS 2024 Camera Ready. First two authors contributed equally. Third and fourth authors contributed equally
专题命中 安全训练 :alignment(title,abstract);RLHF(abstract);分类 cs.CL
Comments The method used in the paper has obvious problems and ambiguities. The security enhancement method we used cannot be considered distillation, but it is described as distillation in the paper, and the experiment lacks comparison and baseline, which has been criticized by many peers. In order to avoid further dissemination, we have decided to withdraw the paper
专题命中 安全训练 :safety(title,abstract);DPO(abstract);分类 cs.CL
Comments 18 pages
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.CL
专题命中 安全训练 :safety(title,abstract);jailbreak(abstract);分类 cs.CL
Comments 17 pages, 8 figures, 2 tables
专题命中 安全训练 :safety(title,abstract);trustworthy(abstract);分类 cs.AI
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.LG
Comments ICML 2024
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.CL
专题命中 安全训练 :safety(title,abstract);harmlessness(abstract);分类 cs.CL
专题命中 安全训练 :safety(title,abstract);AI safety(abstract);分类 cs.CL
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.CL
专题命中 安全训练 :safety(title,abstract);AI safety(abstract);分类 cs.AI
摩擦性政策优化用于大语言模型:认知干预、风险敏感控制与反思对齐
机构 * Brandeis University(布拉德雷大学) ; Colorado State University(科罗拉多州立大学)
专题命中 安全训练 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出摩擦性政策优化框架,通过干预管理认知与规范风险,引入摩擦干预分类和统一方法,提升模型的认知能力与对齐性。
Comments Frictive Policy Optimization; epistemic alignment; risk-sensitive control; LLM alignment; clarification and refusal; preference learning; trust regions; dialogue agents
专题命中 安全训练 :alignment(title,abstract);分类 cs.AI、cs.CY、cs.LG
Comments Pluralistic Alignment Workshop at NeurIPS 2024
DUET:基于同权重分歧的双教师在线策略蒸馏用于禁止合规性
专题命中 安全训练 :DPO(abstract,abstract_cn);alignment(abstract);safety(abstract);分类 cs.CL、cs.LG
AI总结 本研究针对LLM部署中的动态禁止规则合规问题,提出DUET双教师在线策略蒸馏方法,构建工业基准,在Qwen模型上实现高合规性与效用保留,性能优于基线。