arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-24 至 2025-09-24 共收录 8 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 8 篇

2501.04561 2025-09-24 cs.CL cs.CV 79%

OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis

Run Luo, Ting-En Lin, Haonan Zhang, Yuchuan Wu, Xiong Liu, Min Yang, Yongbin Li, Longze Chen, Jiaming Li, Lei Zhang, Xiaobo Xia, Hamid Alinejad-Rokny, Fei Huang

机构 * Shenzhen Key Laboratory for High Performance Data Mining(深圳高性能数据挖掘重点实验室) Shenzhen Institute of Advanced Technology(深圳先进技术研究院) Chinese Academy of Sciences(中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Tongyi Laboratory(通义实验室) University of New South Wales(新南威尔士大学) National University of Singapore(新加坡国立大学) University of Science and Technology of China(中国科学技术大学) MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition(脑启发智能感知与认知重点实验室)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15044 2025-09-24 cs.CL 77%

Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner

Bolian Li, Yanran Wu, Xinyu Luo, Ruqi Zhang

机构 * Department of Computer Science, Purdue University(计算机科学系,普渡大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);safety(abstract);分类 cs.CL

Comments EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24846 2025-09-24 cs.AI cs.CL 73%

MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning

Jingyan Shen, Jiarui Yao, Rui Yang, Yifan Sun, Feng Luo, Rui Pan, Tong Zhang, Han Zhao

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) New York University(纽约大学) Rice University(稻谷大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18632 2025-09-24 cs.CL 70%

A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users

Nishant Balepur, Matthew Shu, Yoo Yeon Sung, Seraphina Goldfarb-Tarrant, Shi Feng, Fumeng Yang, Rachel Rudinger, Jordan Lee Boyd-Graber

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19265 2025-09-24 cs.AI cs.CL 62%

Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World

Saeed Almheiri, Rania Hossam, Mena Attia, Chenxi Wang, Preslav Nakov, Timothy Baldwin, Fajri Koto

机构 * Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 - Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18316 2025-09-24 cs.CL cs.AI 62%

Brittleness and Promise: Knowledge Graph Based Reward Modeling for Diagnostic Reasoning

Saksham Khatwani, He Cheng, Majid Afshar, Dmitriy Dligach, Yanjun Gao

机构 * University of Colorado Boulder(科罗拉多大学博尔德分校) University of Colorado Anschutz(科罗拉多大学安舒茨分校) University of Wisconsin - Madison(威斯康星大学麦迪逊分校) Loyola University Chicago(芝加哥洛约拉大学)

专题命中 偏好对齐 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18661 2025-09-24 cs.IR cs.CL cs.HC 57%

Agentic AutoSurvey: Let LLMs Survey LLMs

Yixin Liu, Yonghui Wu, Denghui Zhang, Lichao Sun

机构 * Lehigh University(莱文大学) University of Florida(佛罗里达大学) Stevens Institute of Technology(史蒂文斯理工学院)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL

Comments 29 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14487 2025-09-24 cs.CV 50%

Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation

Jihao Gu, Yingyao Wang, Meng Cao, Pi Bu, Jun Song, Yancheng He, Shilong Li, Bo Zheng

机构 * Taobao & Tmall Group of Alibaba(淘宝与天猫集团)

专题命中 偏好对齐 :DPO(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏