arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-17 至 2025-11-17 共收录 3 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3 篇

2511.10656 2025-11-17 cs.CL cs.AI 81%

Preference Orchestrator: Prompt-Aware Multi-Objective Alignment for Large Language Models

Biao Liu, Ning Xu, Junming Yang, Xin Geng

机构 * School of Computer Science and Engineering(计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications(新一代人工智能技术及其交叉应用关键实验室) Ministry of Education(教育部)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10686 2025-11-17 cs.CL 77%

A methodological analysis of prompt perturbations and their effect on attack success rates

Tiago Machado, Maysa Malfiza Garcia de Macedo, Rogerio Abreu de Paula, Marcelo Carpinette Grave, Aminat Adebiyi, Luan Soares de Souza, Enrico Santarelli, Claudio Pinhanez

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11370 2025-11-17 cs.IR 50%

SRLF: An Agent-Driven Set-Wise Reflective Learning Framework for Sequential Recommendation

Jiahao Wang, Bokang Fu, Yu Zhu, Yuli Liu

专题命中 偏好对齐 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏