arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-08 至 2025-09-08 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2509.04713 2025-09-08 cs.LG 79%

Natural Spectral Fusion: p-Exponent Cyclic Scheduling and Early Decision-Boundary Alignment in First-Order Optimization

Gongyue Zhang, Honghai Liu

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20309 2025-09-08 cs.CV 78%

Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs

Zitian Wang, Yue Liao, Kang Rong, Fengyun Rao, Yibo Yang, Si Liu

机构 * Beihang University(北航大学) National University of Singapore(国立新加坡大学) King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学)

专题命中 偏好对齐 :alignment(title,abstract)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03730 2025-09-08 cs.AI cs.CL cs.CY cs.LG stat.ML 77%

The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs

Pengrui Han, Rafal Kocielnik, Peiyang Song, Ramit Debnath, Dean Mobbs, Anima Anandkumar, R. Michael Alvarez

机构 * Caltech(加州理工学院) UIUC(伊利诺伊大学) University of Cambridge(剑桥大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.CY

Comments We make public all code and source data at https://github.com/psychology-of-AI/Personality-Illusion for full reproducibility

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05946 2025-09-08 eess.SY cs.SY 50%

InstructMPC: A Human-LLM-in-the-Loop Framework for Context-Aware Control

Ruixiang Wu, Jiahao Ai, Tongxin Li

专题命中 偏好对齐 :DPO(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏