arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-22 至 2025-10-22 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2502.00657 2025-10-22 cs.LG cs.AI cs.CY stat.ML 90%

LLM Safety Alignment is Divergence Estimation in Disguise

Rajdeep Haldar, Ziyi Wang, Qifan Song, Guang Lin, Yue Xing

机构 * Department of Statistics, Purdue University(普渡大学统计系) Department of Statistics, Michigan State University(密歇根州立大学统计系)

专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);RLHF(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10935 2025-10-22 cs.CL 70%

Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documents

Ankan Mullick, Sombit Bose, Rounak Saha, Ayan Kumar Bhowmick, Aditya Vempaty, Prasenjit Dey, Ravi Kokku, Pawan Goyal, Niloy Ganguly

机构 * IIT Kharagpur(印度克达尔普大学) Emergence AI

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.CL

Comments Paper accepted in EMNLP 2025 Main Conference (Full Paper)

Journal ref EMNLP 2025 Main Conference (Full Paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17999 2025-10-22 cs.CY cs.AI cs.HC cs.LG 67%

The Narcissus Hypothesis: Descending to the Rung of Illusion

Riccardo Cadei, Christian Internò

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18849 2025-10-22 cs.CL cs.AI 62%

Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning

Chenghao Zhu, Meiling Tao, Tiannan Wang, Dongyi Ding, Yuchen Eleanor Jiang, Wangchunshu Zhou

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Electronic Science and Technology of China(电子科技大学) South China Agricultural University(华南农业大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

Comments work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18433 2025-10-22 cs.CV cs.AI cs.IR 57%

ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization

Yuanhe Guo, Linxi Xie, Zhuoran Chen, Kangrui Yu, Ryan Po, Guandao Yang, Gordon Wetztein, Hongyi Wen

机构 * NYU(纽约大学) Stanford(斯坦福大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏