arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-17 至 2025-10-17 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 5 篇

2510.01545 2025-10-17 cs.LG cs.AI cs.RO 62%

Predictive Preference Learning from Human Interventions

Haoyuan Cai, Zhenghao Peng, Bolei Zhou

机构 * Department of Computer Science, University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)

专题命中 偏好对齐 :safety(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025 Spotlight. Project page: https://metadriverse.github.io/ppl

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14526 2025-10-17 cs.CV cs.LG 57%

Noise Projection: Closing the Prompt-Agnostic Gap Behind Text-to-Image Misalignment in Diffusion Models

Yunze Tong, Didi Zhu, Zijing Hu, Jinluan Yang, Ziyu Zhao

机构 * Zhejiang University(浙江大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.LG

Comments Appendix will be appended soon

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14200 2025-10-17 cs.CL 57%

RLSR: Reinforcement Learning with Supervised Reward Outperforms SFT in Instruction Following

Zhichao Wang, Andy Wong, Ruslan Belkin

机构 * Inflection AI

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14374 2025-10-17 cs.CV 50%

Spatial Preference Rewarding for MLLMs Spatial Understanding

Han Qiu, Peng Gao, Lewei Lu, Xiaoqin Zhang, Ling Shao, Shijian Lu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) Shanghai AI Laboratory(上海人工智能实验室) Sensetime Research(商汤科技研究院) Zhejiang University of Technology(浙江工业大学) UCAS-Terminus AI Lab,University of Chinese Academy of Sciences(中国科学院大学Terminus AI实验室)

专题命中 偏好对齐 :alignment(abstract)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14257 2025-10-17 cs.IR 50%

Synergistic Integration and Discrepancy Resolution of Contextualized Knowledge for Personalized Recommendation

Lingyu Mu, Hao Deng, Haibo Xing, Kaican Lin, Zhitong Zhu, Yu Zhang, Xiaoyi Zeng, Zhengxiao Liu, Zheng Lin, Jinxin Hu

专题命中 偏好对齐 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏