arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-12 至 2025-09-12 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2509.09055 2025-09-12 cs.CL cs.AI cs.LG 92%

Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M

Piyush Pant

机构 * Saarland University(萨尔兰大学)

专题命中 偏好对齐 :DPO(title,abstract);safety(title,abstract);alignment(abstract);RLHF(abstract)

Comments 17 pages, 3 figures. Code and dataset available at https://github.com/PiyushWithPant/Improving-LLM-Safety-and-Helpfulness-using-SFT-and-DPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20655 2025-09-12 cs.CV cs.CL 83%

Improving Alignment in LVLMs with Debiased Self-Judgment

Sihan Yang, Chenhang Cui, Zihao Zhao, Yiyang Zhou, Weilong Yan, Ying Wei, Huaxiu Yao

机构 * Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学) UNC-Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 偏好对齐 :alignment(title,abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08302 2025-09-12 cs.CL cs.AI 62%

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

Jiahui Li, Lin Li, Tai-wei Chang, Kun Kuang, Long Chen, Jun Zhou, Cheng Yang

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09121 2025-09-12 cs.CL 57%

Compass-v3: Scaling Domain-Specific LLMs for Multilingual E-Commerce in Southeast Asia

Sophia Maria

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07370 2025-09-12 cs.CL 57%

PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions

Yixuan Tang, Yi Yang, Ahmed Abbasi

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) University of Notre Dame(诺丁汉大学)

专题命中 偏好对齐 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏