Noise Contrastive Alignment of Language Models with Explicit Rewards
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.LG
Comments NeurIPS 2024
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.LG
Comments NeurIPS 2024
专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.CL、cs.AI
Comments The first four authors contributed equally, 25 pages
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI
Comments 28 pages
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.CL、cs.AI
Comments 10 pages, 27 figures (including 18 in the appendix), submitted to EMNLP 2024
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI
Comments Equal contribution for the first two authors; To appear in proceedings of the Main Conference on Empirical Methods in Natural Language Processing (EMNLP) 2024
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI
Comments Accepted by EMNLP 2024(Findings)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :safety(title,abstract);jailbreak(abstract);分类 cs.CL、cs.LG
Comments Preprint
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG
Comments Accepted by ACL 2024 Findings
专题命中 偏好对齐 :alignment(title);RLHF(abstract);DPO(abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG
Comments 11 pages, 6 figures. arXiv admin note: text overlap with arXiv:2401.06080
专题命中 偏好对齐 :RLHF(title,abstract);alignment(abstract);分类 cs.CL、cs.LG
Comments 19 pages, 6 figures, 21 tables
专题命中 偏好对齐 :RLHF(title,abstract);jailbreak(abstract);分类 cs.AI、cs.LG
Comments Presented at ICLR 2024
专题命中 偏好对齐 :safety(title,abstract);DPO(abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL、cs.AI
Comments ICLR 2024
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI
Comments Accepted by AAAI 2024
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI
Comments 26 pages, 8 figures. Project website: https://allenai.github.io/re-align/
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI
Comments Accepted by Workshop on Instruction Tuning and Instruction Following at NeurIPS 2023, Submitted to AAAI 2024
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI
Comments Published in NeurIPS 2023 Datasets and Benchmarks
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI、cs.LG
机构 * Peking University(北京大学) ; UC Berkeley(加州大学伯克利分校) ; Center for Human-Compatible AI(人类兼容人工智能中心)
专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY
Comments Journal of Artificial Intelligence Research, in press. Best Paper at NeurIPS 2024 Pluralistic Alignment Workshop