arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-18 至 2025-11-18 共收录 95 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 11 篇

2511.09385 2025-11-18 cs.CL 83%

AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment

Ruibo Deng, Duanyu Feng, Wenqiang Lei

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL

Comments AAAI 2026 AIA oral, our code is available at https://github.com/Shiroha-Offical/AMaPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13007 2025-11-18 cs.AI cs.LG 81%

GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs

Yiyang Zhao, Huiyu Bai, Xuejiao Zhao

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments This paper has been accepted by AAAI 2026-AIA and designated as an oral presentation paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12573 2025-11-18 cs.CL cs.AI 81%

Mitigating Length Bias in RLHF through a Causal Lens

Hyeonji Kim, Sujeong Oh, Sanghack Lee

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06965 2025-11-18 cs.CL cs.AI 81%

Uncovering Factor Level Preferences to Improve Human-Model Alignment

Juhyun Oh, Eunsu Kim, Jiseon Kim, Wenda Xu, Inha Cha, William Yang Wang, Alice Oh

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09724 2025-11-18 cs.AI 79%

UDA: Unsupervised Debiasing Alignment for Pair-wise LLM-as-a-Judge

Yang Zhang, Cunxiang Wang, Lindong Wu, Wenbo Yu, Yidong Wang, Guangsheng Bao, Jie Tang

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13016 2025-11-18 cs.LG 70%

The Good, The Bad, and The Hybrid: A Reward Structure Showdown in Reasoning Models Training

Subramanyam Sahoo

机构 * Berkeley AI Safety Initiative (BASIS) UC Berkeley(伯克利人工智能安全计划(BASIS)加州大学伯克利分校)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.LG

Comments Paper accepted to the 2nd Workshop on Aligning Reinforcement Learning Experimentalists and Theorists (ARLET 2025) at NeurIPS; the paper consists of 14 pages (including the appendix) and contains 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12036 2025-11-18 cs.CE cond-mat.mtrl-sci cs.AI cs.CL cs.LG 67%

Preference Learning from Physics-Based Feedback: Tuning Language Models to Design BCC/B2 Superalloys

Satanu Ghosh, Collin Holgate, Neal R. Brodnik, Doug Downey, Samantha Daly, Tresa M. Pollock, Samuel Carton

机构 * Department of Computer Science, University of New Hampshire(新罕布什尔大学计算机科学系) Materials Department, University of California, Santa Barbara(加州大学圣芭芭拉分校材料系) Department of Mechanical Engineering, University of California, Santa Barbara(加州大学圣芭芭拉分校机械工程系) Allen Institute for Artificial Intelligence(人工智能研究院)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12464 2025-11-18 cs.CL 57%

Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models

Chenglong Wang, Yifu Huo, Yang Gan, Yongyu Mu, Qiaozhi He, Murun Yang, Bei Li, Chunliang Zhang, Tongran Liu, Anxiang Ma, Zhengtao Yu, Jingbo Zhu, Tong Xiao

机构 * School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院) Meituan Inc.(美团公司) NiuTrans Research(牛译研所) CAS Key Laboratory of Behavioral Science, Institute of Psychology, CAS(中国科学院行为科学重点实验室) Kunming University of Science and Technology(昆明理工大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10071 2025-11-18 cs.AI 57%

Advanced Tool Learning and Selection System (ATLASS): A Closed-Loop Framework Using LLM

Mohd Ariful Haque, Justin Williams, Sunzida Siddique, Md. Hujaifa Islam, Hasmot Ali, Kishor Datta Gupta, Roy George

机构 * Clark Atlanta University(克拉克亚特兰大大学) Daffodil International University(达福尔国际大学) Ahsanullah University of Science and Technology(阿沙努拉大学科学与技术学院)

专题命中 偏好对齐 :safety(abstract);分类 cs.AI

Journal ref 2025 IEEE International Conference on Service-Oriented System Engineering (SOSE), Tucson, AZ, USA, 21-24 July 2025, IEEE

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18212 2025-11-18 cs.CL 57%

Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries

Meiling Ning, Zhongbao Zhang, Junda Ye, Jiabao Guo, Qingyuan Guan

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL

Comments After further internal discussion, our author team has decided to withdraw this submission due to the need for several important refinements to the manuscript. All co-authors have been informed and agree with this decision

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07864 2025-11-18 cs.CV 50%

Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization

Tiancheng Yang, Lin Zhang, Jiaye Lin, Guimin Hu, Di Wang, Lijie Hu

机构 * MBZUAI Provable Responsible AI and Data Analytics (PRADA) Lab(可证明负责任的人工智能与数据分析实验室) King Abdullah University of Science and Technology(卡迪夫大学科学与技术大学) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院) University of Copenhagen(哥本哈根大学) Tsinghua University(清华大学)

专题命中 偏好对齐 :DPO(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 7 篇

2511.12982 2025-11-18 cs.CR cs.CV 88%

SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization

Xuankun Rong, Wenke Huang, Tingfeng Wang, Daiguo Zhou, Bo Du, Mang Ye

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) MiLM Plus, Xiaomi Inc.(小米公司)

专题命中 安全训练 :alignment(title,abstract);safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12497 2025-11-18 cs.CL cs.AI cs.CR 84%

SGuard-v1: Safety Guardrail for Large Language Models

JoonHo Lee, HyeonMin Cho, Jaewoong Yun, Hyunjae Lee, JunKyu Lee, Juree Seok

专题命中 安全训练 :safety(title,abstract);AI safety(abstract);分类 cs.CL、cs.AI

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11693 2025-11-18 cs.AI cs.CR cs.CV cs.LG 73%

Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation

Xin Zhao, Xiaojun Chen, Bingshan Liu, Zeyao Liu, Zhendong Zhao, Xiaoyan Gu

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14031 2025-11-18 cs.CL cs.AI cs.LG 71%

Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation

Dongyoon Hahm, Taywon Min, Woogyeol Jin, Kimin Lee

专题命中 安全训练 :safety(abstract,comments);分类 cs.CL、cs.AI、cs.LG;alignment(comments)

Comments Accepted at AAAI 2026 AI Alignment Track, Source code: https://github.com/HahmDY/agentic-ft-safety

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12991 2025-11-18 cs.CL 57%

Fine-Tuned LLMs Know They Don't Know: A Parameter-Efficient Approach to Recovering Honesty

Zeyu Shi, Ziming Wang, Tianyu Chen, Shiqi Gao, Haoyi Zhou, Qingyun Sun, Jianxin Li

专题命中 安全训练 :trustworthy(abstract);分类 cs.CL

Comments Accepted by AAAI 2026 Main Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09073 2025-11-18 cs.FL cs.AI cs.GT 57%

Good-for-MDP State Reduction for Stochastic LTL Planning

Christoph Weinhuber, Giuseppe De Giacomo, Yong Li, Sven Schewe, Qiyi Tang

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 16 pages including appendices, accepted to AAAI 2026; fixed some typoes

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12160 2025-11-18 cs.RO 50%

Game-Theoretic Safe Multi-Agent Motion Planning with Reachability Analysis for Dynamic and Uncertain Environments (Extended Version)

Wenbin Mai, Minghui Liwang, Xinlei Yi, Xiaoyu Xia, Seyyedali Hosseinalipour, Xianbin Wang

机构 * Department of Electrical and Computer Engineering, National University of Singapore(国立新加坡大学电气与计算机工程系) Department of Control Science and Engineering, Shanghai Institute of Intelligent Science and Technology(上海智能科学与技术研究院控制科学与工程系)

专题命中 安全训练 :safety(abstract)

Comments 12 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 11 篇

2511.12217 2025-11-18 cs.LG 83%

AlignTree: Efficient Defense Against LLM Jailbreak Attacks

Gil Goren, Shahar Katz, Lior Wolf

专题命中 越狱攻击 :jailbreak(title);alignment(abstract);safety(abstract);分类 cs.LG

Comments Accepted as an Oral Presentation at the 40th AAAI Conference on Artificial Intelligence (AAAI-26), January 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22564 2025-11-18 cs.CL cs.AI cs.LG 82%

Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs

Xikang Yang, Biyu Zhou, Xuehai Tang, Jizhong Han, Songlin Hu

专题命中 越狱攻击 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06194 2025-11-18 cs.CL 79%

SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation

Lai Jiang, Yuekang Li, Xiaohan Zhang, Youtao Ding, Li Pan

专题命中 越狱攻击 :jailbreak(title,abstract);分类 cs.CL

Comments This paper has been accepted by AAAI 2026 as a poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13548 2025-11-18 cs.CR cs.AI cs.CL 73%

ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models

Siyang Cheng, Gaotian Liu, Rui Mei, Yilin Wang, Kejia Zhang, Kaishuo Wei, Yuqi Yu, Weiping Wen, Xiaojie Wu, Junhua Liu

机构 * iFLYTEK Security Laboratory(iFLYTEK安全实验室) Anhui SparkShield Intelligent Technology Co., Ltd.(安徽SparkShield智能科技有限公司) Peking University(北京大学) School of Automation, University of Electronic Science and Technology of China(电子科技大学自动化学院) Northwest University(西北大学) University of New South Wales(新南威尔士大学) National Computer Network Emergency Response Technical Team/Coordination Center of China(中国国家计算机网络应急技术处置协调中心)

专题命中 越狱攻击 :alignment(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01223 2025-11-18 cs.CR cs.CL 70%

Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge

Ning Xu, Bo Gao, Hui Dou

机构 * School of Computer Science and Technology(计算机科学与技术学院) AnHui University(安徽大学)

专题命中 越狱攻击 :alignment(abstract);jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12782 2025-11-18 cs.CL cs.CR 70%

LLM Reinforcement in Context

Thomas Rivasseau

机构 * McGill University(麦吉尔大学)

专题命中 越狱攻击 :alignment(abstract);jailbreak(abstract);分类 cs.CL

Comments 4 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17089 2025-11-18 cs.CL 70%

Chain-of-Thought Driven Adversarial Scenario Extrapolation for Robust Language Models

Md Rafi Ur Rashid, Vishnu Asutosh Dasu, Ye Wang, Gang Tan, Shagufta Mehnaz

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL

Comments 19 pages, 5 figures. Accepted in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12648 2025-11-18 cs.CR cs.AI cs.LG 62%

Scalable Hierarchical AI-Blockchain Framework for Real-Time Anomaly Detection in Large-Scale Autonomous Vehicle Networks

Rathin Chandra Shit, Sharmila Subudhi

机构 * organization= Dept. of Computer Science \& Engg., International Institute of Information Technology , city= Bhubaneswar , postcode= 751003 , state= Odisha , country= India organization= Dept. of Computer Science, Maharaja Sriram Chandra Bhanja Deo University , city= Baripada , postcode= 757003 , state= Odisha , country= India

专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG

Comments Submitted to the Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12423 2025-11-18 cs.CR cs.LG 57%

GRAPHTEXTACK: A Realistic Black-Box Node Injection Attack on LLM-Enhanced GNNs

Jiaji Ma, Puja Trivedi, Danai Koutra

专题命中 越狱攻击 :prompt injection(abstract);分类 cs.LG

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12149 2025-11-18 cs.CR cs.AI cs.CV 57%

AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models

Jiayu Li, Yunhan Zhao, Xiang Zheng, Zonghuan Xu, Yige Li, Xingjun Ma, Yu-Gang Jiang

机构 * Fudan University(复旦大学) City University of Hong Kong(香港城市大学) Singapore Management University(新加坡管理大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11961 2025-11-18 cs.HC 50%

"Power of Words": Stealthy and Adaptive Private Information Elicitation via LLM Communication Strategies

Shuning Zhang, Jiaqi Bai, Linzhi Wang, Shixuan Li, Xin Yi, Hewu Li

专题命中 越狱攻击 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 2 篇

2511.12295 2025-11-18 cs.CR 78%

Privacy-Preserving Prompt Injection Detection for LLMs Using Federated Learning and Embedding-Based NLP Classification

Hasini Jayathilaka

专题命中 提示注入 :prompt injection(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏