arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-16 至 2025-09-16 共收录 6 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 6 篇

2509.08541 2025-09-16 cs.CL 83%

CM-Align: Consistency-based Multilingual Alignment for Large Language Models

Xue Zhang, Yunlong Liang, Fandong Meng, Songming Zhang, Yufeng Chen, Jinan Xu, Jie Zhou

机构 * Key Laboratory of Big Data & Artificial Intelligence in Transportation, Beijing Jiaotong University, Ministry of Education(大数据与人工智能交通运输 key laboratory,北京交通大学,教育部) School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China(计算机科学与技术学院,北京交通大学,北京,中国) Pattern Recognition Center, WeChat AI, Tencent Inc, China(模式识别中心,微信AI,腾讯公司,中国)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05381 2025-09-16 cs.AI cs.LG 81%

Murphys Laws of AI Alignment: Why the Gap Always Wins

Madhava Gaikwad

机构 * Microsoft(微软)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments Provides a formal impossibility theorem (Murphys Gap) and welcomes collaboration on large-scale experiments and benchmark design

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12301 2025-09-16 cs.LG cs.CL 62%

One Goal, Many Challenges: Robust Preference Optimization Amid Content-Aware and Multi-Source Noise

Amirabbas Afzali, Amirhossein Afsharrad, Seyed Shahabeddin Mousavi, Sanjay Lall

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01062 2025-09-16 cs.LG cs.AI 62%

Offline RLAIF: Piloting VLM Feedback for RL via SFO

Jacob Beck

机构 * Jacob Beck 1

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

Comments Code is provided at https://github.com/jacooba/OfflineRLAIF

Journal ref Published at The RLC 2025 Workshop on Reinforcement Learning Beyond Rewards: Ingredients for Developing Generalist Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11287 2025-09-16 cs.CV cs.CL 57%

Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations

Yifan Lu, Ziqi Zhang, Chunfeng Yuan, Jun Gao, Congxuan Zhang, Xiaojuan Qi, Bing Li, Weiming Hu

机构 * Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information, CASIA(北京多模态信息超级智能安全重点实验室,中国科学院自动化所) State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(多模态人工智能系统国家重点实验室,中国科学院自动化所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Hello Group(Hello集团) Nanchang Hangkong University(南昌航空大学) The University of Hong Kong(香港大学) School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments emnlp 2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00047 2025-09-16 cs.CL 57%

Base Models Beat Aligned Models at Randomness and Creativity

Peter West, Christopher Potts

机构 * Stanford University(斯坦福大学) University of British Columbia(不列颠哥伦比亚大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏