arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-15 至 2025-10-15 共收录 44 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 8 篇

2406.02575 2025-10-15 cs.CL cs.CR cs.LG 90%

Cross-Modal Safety Alignment: Is textual unlearning all you need?

Trishna Chakraborty, Erfan Shayegani, Zikui Cai, Nael Abu-Ghazaleh, M. Salman Asif, Yue Dong, Amit K. Roy-Chowdhury, Chengyu Song

机构 * University of California, Riverside(加州大学河滨分校)

专题命中 偏好对齐 :alignment(title,abstract);safety(title,abstract);RLHF(abstract);分类 cs.CL、cs.LG

Comments Accepted by EMNLP 2024 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12044 2025-10-15 cs.CL cs.AI 84%

Hierarchical Alignment: Surgical Fine-Tuning via Functional Layer Specialization in Large Language Models

Yukun Zhang, Qi Dong

机构 * The Chinese University of Hong Kong(香港中文大学) Fudan University(复旦大学)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12195 2025-10-15 cs.CL 83%

DPO-Tuned Large Language Models for Segmentation in Simultaneous Speech Translation

Zeyu Yang, Satoshi Nakamura

专题命中 偏好对齐 :DPO(title,abstract);alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23564 2025-10-15 cs.AI cs.CL 81%

Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment

Samuel Yeh, Sharon Li

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23223 2025-10-15 cs.LG cs.AI cs.CL cs.GT 75%

COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences

Yixin Liu, Argyris Oikonomou, Weiqiang Zheng, Yang Cai, Arman Cohan

机构 * Yale University(耶鲁大学) Allen Institute for AI(人工智能研究院)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11978 2025-10-15 cs.LG cs.AI 73%

Learning Dynamics of VLM Finetuning

Jusheng Zhang, Kaitong Cai, Jing Yang, Keze Wang

机构 * Sun Yat-sen University(中山大学) X-Era AI Lab(X-Era人工智能实验室)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12014 2025-10-15 cs.IR cs.LG 57%

Embedding the Teacher: Distilling vLLM Preferences for Scalable Image Retrieval

Eric He, Akash Gupta, Adian Liusie, Vatsal Raina, Piotr Molenda, Shirom Chabra, Vyas Raina

机构 * University of Cambridge(剑桥大学) Apta

专题命中 偏好对齐 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26041 2025-10-15 cs.CL 57%

Unspoken Hints: Accuracy Without Acknowledgement in LLM Reasoning

Arash Marioriyad, Shaygan Adim, Nima Alighardashi, Mahdieh Soleymani Banghshah, Mohammad Hossein Rohban

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL

Comments 5 Pages, 4 Figures, 4 Tables

Journal ref 39th Conference on Neural Information Processing Systems, 2025, Workshop: Reliable ML from Unreliable Data

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 6 篇

2509.01909 2025-10-15 cs.AI cs.CL cs.CY cs.HC cs.SC 90%

Oyster-I: Beyond Refusal -- Constructive Safety Alignment for Responsible Language Models

Ranjie Duan, Jiexi Liu, Xiaojun Jia, Shiji Zhao, Ruoxi Cheng, Fengxiang Wang, Cheng Wei, Yong Xie, Chang Liu, Defeng Li, Yinpeng Dong, Yichi Zhang, Yuefeng Chen, Chongwen Wang, Xingjun Ma, Xingxing Wei, Yang Liu, Hang Su, Jun Zhu, Xinfeng Li, Yitong Sun, Jie Zhang, Jinzhao Hu, Sha Xu, Wenchao Yang, Yitong Yang, Xingyao Zhang, Yingshui Tan, Jialing Tao, Hui Xue

机构 * Alibaba AAIG(阿里巴巴AAIG)

专题命中 安全训练 :alignment(title,abstract);safety(title,abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Technical Report Code & Model weights available: https://github.com/Alibaba-AAIG/Oyster

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12083 2025-10-15 cs.CL cs.AI 81%

An AI-Based Behavioral Health Safety Filter and Dataset for Identifying Mental Health Crises in Text-Based Conversations

Benjamin W. Nelson, Celeste Wong, Matthew T. Silvestrini, Sooyoon Shin, Alanna Robinson, Jessica Lee, Eric Yang, John Torous, Andrew Trister

专题命中 安全训练 :safety(title,abstract);分类 cs.CL、cs.AI

Comments Main Text: 2943; Abstract: 256; Tables and Figures: 5

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07961 2025-10-15 cs.RO cs.AI 79%

BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models

Peiyan Li, Yixiang Chen, Hongtao Wu, Xiao Ma, Xiangnan Wu, Yan Huang, Liang Wang, Tao Kong, Tieniu Tan

机构 * CASIA(中国科学院自动化研究所) ByteDance Seed(字节跳动种子实验室) UCAS(中国科学院大学) FiveAges NJU(南京大学)

专题命中 安全训练 :alignment(title,abstract);分类 cs.AI

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23441 2025-10-15 cs.CL 70%

Cognition-of-Thought Elicits Social-Aligned Reasoning in Large Language Models

Xuanming Zhang, Yuxuan Chen, Samuel Yeh, Sharon Li

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Tsinghua University(清华大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12428 2025-10-15 cs.AI 57%

Biased-Attention Guided Risk Prediction for Safe Decision-Making at Unsignalized Intersections

Chengyang Dong, Nan Guo

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12477 2025-10-15 cs.RO 50%

A Task-Efficient Reinforcement Learning Task-Motion Planner for Safe Human-Robot Cooperation

Gaoyuan Liu, Joris de Winter, Kelly Merckaert, Denis Steckelmacher, Ann Nowe, Bram Vanderborght

机构 * Department of Mechanical Engineering, Vrije Universiteit Brussel(布鲁塞尔自由大学机械工程系) imec Flanders Make(弗拉芒制造) Artificial Intelligence (AI) Lab, Vrije Universiteit Brussel(布鲁塞尔自由大学人工智能实验室)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 3 篇

2507.07146 2025-10-15 cs.LG cs.CL 84%

Attention-Aware GNN-based Input Defense against Multi-Turn LLM Jailbreak

Zixuan Huang, Kecheng Huang, Lihao Yin, Bowei He, Huiling Zhen, Mingxuan Yuan, Zili Shao

机构 * The Chinese University of Hong Kong(香港中文大学) Noah’s Ark Lab, Huawei(华为诺亚实验室) City University of Hong Kong(香港城市大学)

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10601 2025-10-15 cs.CL cs.AI 62%

When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers

Divij Handa, Zehua Zhang, Amir Saeidi, Shrinidhi Kumbhar, Md Nayem Uddin, Aswin RRV, Chitta Baral

机构 * Arizona State University(亚利桑那州立大学)

专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI

Comments Published in Reliable ML from Unreliable Data workshop @ NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11837 2025-10-15 cs.CR cs.AI 61%

Countermind: A Multi-Layered Security Architecture for Large Language Models

Dominik Schwarz

机构 * Independent Researcher(独立研究者)

专题命中 越狱攻击 :prompt injection(abstract,comments);分类 cs.AI

Comments 33 pages, 3 figures, 6 tables. Keywords: LLM security; defense-in-depth; prompt injection; activation steering; multimodal sandbox; threat modeling

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 红队测试 1 篇

2510.11823 2025-10-15 cs.CR cs.AI 83%

BlackIce: A Containerized Red Teaming Toolkit for AI Security Testing

Caelin Kaplan, Alexander Warnecke, Neil Archibald

机构 * Databricks

专题命中 红队测试 :red teaming(title,abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 1 篇

2508.00378 2025-10-15 cs.AI cs.CV 57%

CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding

Shixin Yi, Lin Shang

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments The paper is not yet mature and needs further improvement

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 隐私与版权 1 篇

2510.12727 2025-10-15 cs.LG cs.AI cs.DC 62%

Hierarchical Federated Learning for Crop Yield Prediction in Smart Agricultural Production Systems

Anas Abouaomar, Mohammed El hanjri, Abdellatif Kobbane, Anis Laouiti, Khalid Nafil

机构 * ENSIAS, Mohammed V University in Rabat(ENSIAS,摩洛哥拉巴特穆罕默德五世大学) Samovar, Télécom SudParis, Institut Polytechnique de Paris(Samovar,电信南巴黎,巴黎理工学院)

专题命中 隐私与版权 :alignment(abstract);分类 cs.AI、cs.LG

Comments 6 pages, 3 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 9 篇

2502.00757 2025-10-15 cs.CR cs.AI cs.NE 86%

AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-Improvement

J Rosser, Jakob Foerster

机构 * University of Oxford(牛津大学) FLAIR University of Oxford(牛津大学FLAIR)

专题命中 安全评测 :safety(title,abstract);AI safety(title);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12133 2025-10-15 cs.CL cs.AI 84%

SafeMT: Multi-turn Safety for Multimodal Language Models

Han Zhu, Juntao Dai, Jiaming Ji, Haoran Li, Chengkun Cai, Pengcheng Wen, Chi-Min Chan, Boyuan Chen, Yaodong Yang, Sirui Han, Yike Guo

机构 * Hong Kong University of Science and Technology(香港理工大学) Peking University(北京大学) University of Edinburgh(爱丁堡大学)

专题命中 安全评测 :safety(title,abstract);jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12713 2025-10-15 cs.AI 57%

Towards Robust Artificial Intelligence: Self-Supervised Learning Approach for Out-of-Distribution Detection

Wissam Salhab, Darine Ameyed, Hamid Mcheick, Fehmi Jaafar

机构 * University of Quebec at Chicoutimi(魁北克大学恰普蒂米分校)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12704 2025-10-15 cs.CV cs.AI 57%

Hybrid Explanation-Guided Learning for Transformer-Based Chest X-Ray Diagnosis

Shelley Zixin Shu, Haozhe Luo, Alexander Poellinger, Mauricio Reyes

机构 * ARTORG Center for Biomedical Engineering Research, University of Bern(ARTORG生物医学工程研究中心,伯恩大学) Inselspital (Bern University Hospital)(Inselspital(伯恩大学医院)) Insel Gruppe Bern Universitätsinstitut für Diagnostische, Interventionelle und Pädiatrische Radiologie(Bern大学诊断、介入和儿科放射学研究所) Department of Radiation Oncology, Inselspital, Bern University Hospital(放射肿瘤科,Inselspital,伯恩大学医院)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted by iMIMIC at MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12316 2025-10-15 cs.CL 57%

Beating Harmful Stereotypes Through Facts: RAG-based Counter-speech Generation

Greta Damo, Elena Cabrio, Serena Villata

机构 * Université Côte d’Azur, CNRS, Inria, I3S, France(法国大学-科蒂-阿祖尔大学、国家科学研究中心、法国国家信息与自动化研究所、I3S研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12287 2025-10-15 cs.CV cs.CL 57%

Vision Language Models Map Logos to Text via Semantic Entanglement in the Visual Projector

Sifan Li, Hongkai Chen, Yujun Cai, Qingwen Ye, Liyang Chen, Junsong Yuan, Yiwei Wang

机构 * University of California, Merced(加州大学梅尔德分校) vivo Mobile Communication Co., Ltd.(vivo移动通信有限公司) University of Queensland(昆士兰大学) UCLA(加州大学洛杉矶分校) University at Buffalo(布法罗大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12224 2025-10-15 cs.AI 57%

MedKGEval: A Knowledge Graph-Based Multi-Turn Evaluation Framework for Open-Ended Patient Interactions with Clinical LLMs

Yuechun Yu, Han Ying, Haoan Jin, Wenjian Jiang, Dong Xian, Binghao Wang, Zhou Yang, Mengyue Wu

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11745 2025-10-15 cs.LG 57%

Think as a Doctor: An Interpretable AI Approach for ICU Mortality Prediction

Qingwen Li, Xiaohang Zhao, Xiao Han, Hailiang Huang, Lanjuan Liu

机构 * School of Information Management & Engineering, Shanghai University of Finance and Economics(上海金融学院信息管理与工程学院) Key Laboratory of Data Intelligence and Management (Beihang University), Ministry of Industry and Information Technology, School of Economics and Management, Beihang University(北京航空航天大学数据智能与管理重点实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 42 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12190 2025-10-15 cs.CV 50%

Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos

Shingo Yokoi, Kento Sasaki, Yu Yamaguchi

机构 * Turing Inc.(图灵公司)

专题命中 安全评测 :safety(abstract)

Comments 2nd Place Winner, ICCV 2025 2COOOL Competition

详情

展开后加载摘要…

URL PDF HTML 收藏

8. AI治理与伦理 1 篇

2505.18943 2025-10-15 cs.CL 57%

MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems

Xuanming Zhang, Yuxuan Chen, Samuel Yeh, Sharon Li

机构 * Tsinghua University(清华大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

Comments NeurIPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏