arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-11-13 至 2025-11-13 共收录 42 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 7 篇

2511.09105 2025-11-13 cs.LG cs.AI 87%

Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment

Shigeki Kusaka, Keita Saito, Mikoto Kudo, Takumi Tanabe, Akifumi Wachi, Youhei Akimoto

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.AI、cs.LG

Comments accepted for AAAI 2026 Special Track on AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08594 2025-11-13 cs.CL 85%

Diverse Preference Learning for Capabilities and Alignment

Stewart Slocum, Asher Parker-Sartori, Dylan Hadfield-Menell

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);DPO(abstract);分类 cs.CL

Journal ref 13th International Conference on Learning Representations (ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.17417 2025-11-13 cs.CV 85%

Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment

Robert Wijaya, Ngoc-Bao Nguyen, Ngai-Man Cheung

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13264 2025-11-13 cs.CL 79%

OpenGenAlign: A Preference Dataset and Benchmark for Trustworthy Reward Modeling in Open-Ended, Long-Context Generation

Hanning Zhang, Juntong Song, Juno Zhu, Yuanhao Wu, Tong Zhang, Cheng Niu

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) NewsBreak

专题命中 偏好对齐 :trustworthy(title);safety(abstract);分类 cs.CL

Comments Preprint update

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08835 2025-11-13 cs.CL cs.AI 62%

Beyond Task-Oriented and Chitchat Dialogues: Proactive and Transition-Aware Conversational Agents

Yejin Yoon, Yuri Son, Namyoung So, Minseo Kim, Minsoo Cho, Chanhee Park, Seungshin Lee, Taeuk Kim

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments accepted to EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23658 2025-11-13 cs.LG cs.AI 62%

Aligning Diffusion Language Models via Unpaired Preference Optimization

Vaibhav Jindal, Hejian Sang, Chun-Mao Lai, Yanning Chen, Zhipeng Wang

机构 * LinkedIn Corporation(LinkedIn公司) University of California San Diego(加州大学圣地亚哥分校)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02087 2025-11-13 cs.CL 57%

When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models

Keyu Wang, Jin Li, Shu Yang, Zhuoran Zhang, Di Wang

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 2 篇

2511.08631 2025-11-13 cs.CY 88%

Enabling Frontier Lab Collaboration to Mitigate AI Safety Risks

Nicholas Felstead

专题命中 安全训练 :safety(title,abstract);AI safety(title,abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09013 2025-11-13 cs.RO cs.CV 50%

UniMM-V2X: MoE-Enhanced Multi-Level Fusion for End-to-End Cooperative Autonomous Driving

Ziyi Song, Chen Xia, Chenbing Wang, Haibao Yu, Sheng Zhou, Zhisheng Niu

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 2 篇

2511.08597 2025-11-13 cs.CL cs.AI 62%

Self-HarmLLM: Can Large Language Model Harm Itself?

Heehwan Kim, Sungjune Park, Daeseon Choi

机构 * Soongsil University(首尔大学)

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06151 2025-11-13 cs.CR cs.AI 57%

Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems

Haowei Wang, Rupeng Zhang, Junjie Wang, Mingyang Li, Yuekai Huang, Dandan Wang, Qing Wang

专题命中 越狱攻击 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 1 篇

2503.01908 2025-11-13 cs.CR cs.AI cs.LG 81%

UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning

Jiawei Zhang, Shuang Yang, Bo Li

机构 * Department of Computer Science, University of Chicago(芝加哥大学计算机科学系) Meta Department of Computer Science, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系)

专题命中 红队测试 :red teaming(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 提示注入 3 篇

2511.01634 2025-11-13 cs.CR cs.AI 85%

Prompt Injection as an Emerging Threat: Evaluating the Resilience of Large Language Models

Daniyal Ganiuly, Assel Smaiyl

专题命中 提示注入 :prompt injection(title,abstract);alignment(abstract);safety(abstract);分类 cs.AI

Comments 10 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12815 2025-11-13 cs.CR cs.AI cs.CL cs.LG 82%

Formalizing and Benchmarking Prompt Injection Attacks and Defenses

Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, Neil Zhenqiang Gong

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Published in USENIX Security Symposium 2024; the model sizes for closed-source models are from blog posts. For slides, see https://people.duke.edu/~zg70/code/PromptInjection.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11358 2025-11-13 cs.CR cs.AI 79%

DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks

Yupei Liu, Yuqi Jia, Jinyuan Jia, Dawn Song, Neil Zhenqiang Gong

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments Distinguished Paper Award in IEEE Symposium on Security and Privacy, 2025. For slides, see https://people.duke.edu/~zg70/code/PromptInjection.pdf

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 幻觉与事实性 1 篇

2509.10227 2025-11-13 cs.LG physics.app-ph 57%

A Certifiable Machine Learning-Based Pipeline to Predict Fatigue Life of Aircraft Structures

Ángel Ladrón, Miguel Sánchez-Domínguez, Javier Rozalén, Fernando R. Sánchez, Javier de Vicente, Lucas Lacasa, Eusebio Valero, Gonzalo Rubio

机构 * ETSIAE-UPM - School of Aeronautics, Universidad Politécnica de Madrid(西班牙马德里理工大学航空学院) Institute for Cross-Disciplinary Physics and Complex Systems (IFISC), CSIC-UIB(跨学科物理与复杂系统研究所(IFISC)) Center for Computational Simulation, Universidad Politécnica de Madrid, Campus de Montegancedo, Boadilla del Monte(马德里理工大学计算模拟中心)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG

Comments 34 pages, 17 figures

Journal ref Engineering Failure Analysis, Volume 184, 2026, 110334

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 13 篇

2511.04962 2025-11-13 cs.CL cs.AI 73%

Too Good to be Bad: On the Failure of LLMs to Role-Play Villains

Zihao Yi, Qingxuan Jiang, Ruotian Ma, Xingyu Chen, Qu Yang, Mengru Wang, Fanghua Ye, Ying Shen, Zhaopeng Tu, Xiaolong Li, Linus

机构 * Tencent(腾讯)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08702 2025-11-13 cs.LG cs.AI cs.CR cs.CY 67%

FAIRPLAI: A Human-in-the-Loop Approach to Fair and Private Machine Learning

David Sanchez, Holly Lopez, Michelle Buraczyk, Anantaa Kotal

机构 * Dept. of Computer Science, The University of Texas at El Paso(得克萨斯大学埃尔帕索分校计算机科学系) Dept. of Mathematics, Mountain View High School(山景高中数学系) Dept. of Mathematics, El Paso Independent School District(埃尔帕索独立学区数学系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09067 2025-11-13 cs.CL cs.AI 62%

MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique

Gailun Zeng, Ziyang Luo, Hongzhan Lin, Yuchen Tian, Kaixin Li, Ziyang Gong, Jianxiong Guo, Jing Ma

机构 * Hong Kong Baptist University(香港 Baptist 大学) Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist 大学) National University of Singapore(新加坡国立大学) Beijing Normal University(北京师范大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments 28 pages, 14 figures, 19 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09363 2025-11-13 cs.AI cs.GT cs.LG 62%

ElicitationGPT: Text Elicitation Mechanisms via Language Models

Yifan Wu, Jason Hartline

机构 * Microsoft Research(微软研究院) Northwestern University(西北大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09443 2025-11-13 cs.CV cs.AI 57%

BronchOpt : Vision-Based Pose Optimization with Fine-Tuned Foundation Models for Accurate Bronchoscopy Navigation

Hongchao Shu, Roger D. Soberanis-Mukul, Jiru Xu, Hao Ding, Morgan Ringel, Mali Shen, Saif Iftekar Sayed, Hedyeh Rafii-Tari, Mathias Unberath

机构 * Johnson & Johnson MedTech(强生医疗科技)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03133 2025-11-13 cs.CL 57%

ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models

Boyang Xue, Qi Zhu, Rui Wang, Sheng Wang, Hongru Wang, Minda Hu, Fei Mi, Yasheng Wang, Lifeng Shang, Qun Liu, Kam-Fai Wong

机构 * The Chinese University of Hong Kong(香港中文大学) Huawei Noah’s Ark Lab(华为诺亚实验室) The University of Hong Kong(香港大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23015 2025-11-13 cs.CL 57%

Detecting Stealthy Backdoor Samples based on Intra-class Distance for Large Language Models

Jinwen Chen, Hainan Zhang, Fei Sun, Qinnan Zhang, Sijia Wen, Ziwei Wang, Zhiming Zheng

机构 * Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(未来区块链与隐私计算先进创新中心) Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments EMNLP2025Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16774 2025-11-13 cs.CL 57%

IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models

Yiming Gao, Bin Wang, Chengwei Wei, Shuo Sun, AiTi Aw

机构 * Nanyang Technological University (NTU)(南洋理工大学) MiroMind(米罗Mind) Institute for Infocomm Research (I 2 R)(信息与通信研究院) A*STAR(科技研究局)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Link: https://github.com/AudioLLMs/AudioBench/tree/main/IFEval-Audio

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06227 2025-11-13 cs.CL 57%

LExT: Towards Evaluating Trustworthiness of Natural Language Explanations

Krithi Shailya, Shreya Rajpal, Gokul S Krishnan, Balaraman Ravindran

机构 * Centre for Responsible AI, IIT Madras(责任人工智能中心,IIT马德拉斯)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08649 2025-11-13 q-bio.QM cs.AI 57%

Bio AI Agent: A Multi-Agent Artificial Intelligence System for Autonomous CAR-T Cell Therapy Development with Integrated Target Discovery, Toxicity Prediction, and Rational Molecular Design

Yi Ni, Liwei Zhu, Shuai Li

机构 * Bio LIMS INC(Bio LIMS公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 12 pages, 0 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15153 2025-11-13 cs.CL 57%

Evaluating Deep Unlearning in Large Language Models

Ruihan Wu, Chhavi Yadav, Russ Salakhutdinov, Kamalika Chaudhuri

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02557 2025-11-13 eess.IV cs.CV 50%

RL-U$^2$Net: A Dual-Branch UNet with Reinforcement Learning-Assisted Multimodal Feature Fusion for Accurate 3D Whole-Heart Segmentation

Jierui Qu, Jianchun Zhao

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15526 2025-11-13 eess.IV cs.CV 50%

Multi-scale Cascaded Foundation Model for Whole-body Organs-at-risk Segmentation

Rui Hao, Dayu Tan, Qiankun Li, Chunhou Zheng, Weimin Zhong, Zhigang Zeng

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学) Institute of Artificial Intelligence, Huazhong University of Science and Technology(人工智能研究院,华中科技大学) Hubei Key Laboratory of Brain-Inspired Intelligent Systems, Huazhong University of Science and Technology(湖北省脑启发智能系统重点实验室,华中科技大学) Key Laboratory of Image Processing and Intelligent Control (Huazhong University of Science and Technology), Ministry of Education(图像处理与智能控制重点实验室(华中科技大学),教育部) Key Laboratory of Intelligent Computing and Signal Processing, Ministry of Education, Anhui University(智能计算与信号处理重点实验室(安徽大学),教育部) College of Computing and Data Science (CCDS), Nanyang Technological University(计算与数据科学学院(CCDS),南洋理工大学) East China University of Science and Technology(东华大学)

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

8. AI治理与伦理 3 篇

2509.10590 2025-11-13 cs.CY cs.AI 66%

Machine Unlearning for Responsible and Adaptive AI in Education

Betty Mayeku, Sandra Hummel, Parisa Memarmoshrefi

机构 * Leipzig University(莱比锡大学) Technical University Dresden(德累斯顿技术大学) University of Göttingen(哥廷根大学)

专题命中 AI治理与伦理 :trustworthy(abstract,comments);分类 cs.AI、cs.CY

Comments Accepted paper - ESORICS 2025 - International Workshop on Secure and Trustworthy Machine Unlearning Systems (STMUS)

详情

展开后加载摘要…

URL PDF HTML 收藏