arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-20 至 2025-10-20 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2510.10556 2025-10-20 cs.IR 78%

Self-Supervised Representation Learning with ID-Content Modality Alignment for Sequential Recommendation

Donglin Zhou, Weike Pan, Zhong Ming

专题命中 偏好对齐 :alignment(title,abstract)

Comments The article has been accepted by Frontiers of Computer Science (FCS), with the DOI: {10.1007/s11704-025-50269-4}

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15716 2025-10-20 cs.AI 77%

Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences

Keertana Chidambaram, Karthik Vinary Seetharaman, Vasilis Syrgkanis

机构 * Stanford University(斯坦福大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18492 2025-10-20 cs.CR cs.AI cs.LG 73%

GuardReasoner: Towards Reasoning-based LLM Safeguards

Yue Liu, Hongcheng Gao, Shengfang Zhai, Yufei He, Jun Xia, Zhengyu Hu, Yulin Chen, Xihong Yang, Jiaheng Zhang, Stan Z. Li, Hui Xiong, Bryan Hooi

机构 * NUS(新加坡国立大学) HKUST (Guangzhou)(香港科技大学(广州)) Westlake University(西湖大学)

专题命中 偏好对齐 :DPO(abstract);safety(abstract);分类 cs.AI、cs.LG

Comments 22 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11620 2025-10-20 cs.CL 57%

Enhancing Long Chain-of-Thought Reasoning through Multi-Path Plan Aggregation

Siheng Xiong, Ali Payani, Faramarz Fekri

机构 * Georgia Institute of Technology(佐治亚理工学院) Cisco Research(思科研究)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏