arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-13 至 2025-08-13 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2508.08509 2025-08-13 cs.CL cs.AI 86%

Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression

Jadie Adams, Brian Hu, Emily Veenhuis, David Joy, Bharadwaj Ravichandran, Aaron Bray, Anthony Hoogs, Arslan Basharat

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);harmlessness(abstract);分类 cs.CL、cs.AI

Comments AIES '25: Proceedings of the 2025 AAAI/ACM Conference on AI, Ethics, and Society

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08466 2025-08-13 cs.CL 83%

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints

Daren Yao, Jinsong Yuan, Ruike Chen

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08550 2025-08-13 cs.SD cs.CL 79%

Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization

Chaoqun Cui, Liangbin Huang, Shijing Wang, Zhe Tong, Zhaolong Huang, Xiao Zeng, Xiaofeng Liu

机构 * Alibaba Digital Media and Entertainment Group(阿里巴巴数字媒体与娱乐集团) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence, Beijing Jiaotong University(北京交通大学交通数据挖掘与具身智能重点实验室)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

Comments This paper is accepted by ACL2025 (Main)

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025: 4524-4546

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05496 2025-08-13 cs.AI 70%

InfiAlign: A Scalable and Sample-Efficient Framework for Aligning LLMs to Enhance Reasoning Capabilities

Shuo Cai, Su Lu, Qi Zhou, Kejing Yang, Zhijie Sang, Congkai Xie, Hongxia Yang

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏