arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-16 至 2025-09-16 共收录 61 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 6 篇

2509.08541 2025-09-16 cs.CL 83%

CM-Align: Consistency-based Multilingual Alignment for Large Language Models

Xue Zhang, Yunlong Liang, Fandong Meng, Songming Zhang, Yufeng Chen, Jinan Xu, Jie Zhou

机构 * Key Laboratory of Big Data & Artificial Intelligence in Transportation, Beijing Jiaotong University, Ministry of Education(大数据与人工智能交通运输 key laboratory,北京交通大学,教育部) School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China(计算机科学与技术学院,北京交通大学,北京,中国) Pattern Recognition Center, WeChat AI, Tencent Inc, China(模式识别中心,微信AI,腾讯公司,中国)

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05381 2025-09-16 cs.AI cs.LG 81%

Murphys Laws of AI Alignment: Why the Gap Always Wins

Madhava Gaikwad

机构 * Microsoft(微软)

专题命中 偏好对齐 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments Provides a formal impossibility theorem (Murphys Gap) and welcomes collaboration on large-scale experiments and benchmark design

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12301 2025-09-16 cs.LG cs.CL 62%

One Goal, Many Challenges: Robust Preference Optimization Amid Content-Aware and Multi-Source Noise

Amirabbas Afzali, Amirhossein Afsharrad, Seyed Shahabeddin Mousavi, Sanjay Lall

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01062 2025-09-16 cs.LG cs.AI 62%

Offline RLAIF: Piloting VLM Feedback for RL via SFO

Jacob Beck

机构 * Jacob Beck 1

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

Comments Code is provided at https://github.com/jacooba/OfflineRLAIF

Journal ref Published at The RLC 2025 Workshop on Reinforcement Learning Beyond Rewards: Ingredients for Developing Generalist Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11287 2025-09-16 cs.CV cs.CL 57%

Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations

Yifan Lu, Ziqi Zhang, Chunfeng Yuan, Jun Gao, Congxuan Zhang, Xiaojuan Qi, Bing Li, Weiming Hu

机构 * Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information, CASIA(北京多模态信息超级智能安全重点实验室,中国科学院自动化所) State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(多模态人工智能系统国家重点实验室,中国科学院自动化所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Hello Group(Hello集团) Nanchang Hangkong University(南昌航空大学) The University of Hong Kong(香港大学) School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments emnlp 2025 accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00047 2025-09-16 cs.CL 57%

Base Models Beat Aligned Models at Randomness and Creativity

Peter West, Christopher Potts

机构 * Stanford University(斯坦福大学) University of British Columbia(不列颠哥伦比亚大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 7 篇

2504.16980 2025-09-16 cs.LG 83%

Safety Pretraining: Toward the Next Generation of Safe AI

Pratyush Maini, Sachin Goyal, Dylan Sam, Alex Robey, Yash Savani, Yiding Jiang, Andy Zou, Matt Fredrikson, Zacharcy C. Lipton, J. Zico Kolter

专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08486 2025-09-16 cs.CL 77%

Too Helpful, Too Harmless, Too Honest or Just Right?

Gautam Siddharth Kashyap, Mark Dras, Usman Naseem

机构 * School of Computing, Macquarie University(计算机学院,麦考瑞大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);harmlessness(abstract);分类 cs.CL

Comments EMNLP'25 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11431 2025-09-16 cs.AI cs.CL 62%

Securing AI Agents: Implementing Role-Based Access Control for Industrial Applications

Aadil Gani Ganie

专题命中 安全训练 :prompt injection(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02465 2025-09-16 cs.CL cs.AI 62%

Revealing the Inherent Instructability of Pre-Trained Language Models

Seokhyun An, Minji Kim, Hyounghun Kim

机构 * Department of Computer Science and Engineering, UNIST(UNIST计算机科学与工程系) Graduate School of Artificial Intelligence, POSTECH(POSTECH人工智能研究生院) Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系)

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments Findings of EMNLP 2025 (32 pages). Code available at https://github.com/seokhyunan/response-tuning

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15776 2025-09-16 cs.CL cs.IR 57%

ConvSearch-R1: Enhancing Query Reformulation for Conversational Search with Reasoning via Reinforcement Learning

Changtai Zhu, Siyin Wang, Ruijun Feng, Kai Song, Xipeng Qiu

机构 * Fudan University(复旦大学) ByteDance Inc(字节跳动公司) University of New South Wales(新南威尔士大学)

专题命中 安全训练 :alignment(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 at the Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10478 2025-09-16 cs.NI cs.LG cs.SY eess.SY 57%

The LLM as a Network Operator: A Vision for Generative AI in the 6G Radio Access Network

Oluwaseyi Giwa, Michael Adewole, Tobi Awodumila, Pelumi Aderinto

机构 * African Institute for Mathematical Sciences(非洲数学科学研究所)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Submitted to Workshop on AI and ML for Next-Generation Wireless Communications and Networking, NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12085 2025-09-16 eess.SY cs.SY 50%

Compositional shield synthesis for safe reinforcement learning in partial observability

Steven Carr, Georgios Bakirtzis, Ufuk Topcu

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 5 篇

2508.16347 2025-09-16 cs.CR cs.AI 85%

Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs

Yu Yan, Sheng Sun, Zhe Wang, Yijun Lin, Zenghao Duan, zhifei zheng, Min Liu, Zhiyi yin, Jianping Zhang

机构 * State Key Lab of Processors, Institute of Computing Technology, CAS(中国科学院计算技术研究所状态关键实验室) University of Chinese Academy of Sciences(中国科学院大学) People’s Public Security University of China(中国人民公安大学) Chinese University of Hong Kong(香港大学)

专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11835 2025-09-16 cs.CL cs.AI 73%

Multilingual Collaborative Defense for Large Language Models

Hongliang Li, Jinan Xu, Gengping Cui, Changhao Guan, Fengran Mo, Kaiyu Huang

机构 * Key Laboratory of Big Data & Artificial Intelligence in Transportation (Beijing Jiaotong University), Ministry of Education(大数据与人工智能交通运输 key laboratory(北京交通大学)) School of Computer Science and Technology, Beijing Jiaotong University(计算机科学与技术学院(北京交通大学)) University of Montreal(蒙特利尔大学)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

Comments 21 pages, 4figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11141 2025-09-16 cs.CL 70%

When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity

Shiyao Cui, Xijia Feng, Yingkang Wang, Junxiao Yang, Zhexin Zhang, Biplab Sikdar, Hongning Wang, Han Qiu, Minlie Huang

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10931 2025-09-16 cs.AI cs.CL 62%

Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding

Seongho Joo, Hyukhun Koh, Kyomin Jung

专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10563 2025-09-16 cs.CR 50%

Enhancing IoMT Security with Explainable Machine Learning: A Case Study on the CICIOMT2024 Dataset

Mohammed Yacoubi, Omar Moussaoui, C. Drocourt

专题命中 越狱攻击 :safety(abstract)

Journal ref The Third Edition of the International Conference on Connected Objects and Artificial Intelligence (COCIA'2025), Apr 2025, Casablanca (Maroc), Morocco

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 提示注入 3 篇

2410.14827 2025-09-16 cs.CR cs.AI cs.CL cs.LG 89%

Enhancing Prompt Injection Attacks to LLMs via Poisoning Alignment

Zedian Shao, Hongbin Liu, Jaden Mu, Neil Zhenqiang Gong

机构 * Georgia Institute of Technology(佐治亚理工学院) Duke University(杜克大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 提示注入 :alignment(title,abstract);prompt injection(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10540 2025-09-16 cs.CR cs.AI 79%

EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System

Pavan Reddy, Aditya Sanjay Gujral

专题命中 提示注入 :prompt injection(title,abstract);分类 cs.AI

Comments 8 pages content, 1 page references, 2 figures, Published at AAAI Fall Symposium Series 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12188 2025-09-16 cs.CR cs.LG 57%

Multi-Agent Systems Execute Arbitrary Malicious Code

Harold Triedman, Rishi Jha, Vitaly Shmatikov

机构 * Cornell Tech(康奈尔科技)

专题命中 提示注入 :prompt injection(abstract);分类 cs.LG

Comments 33 pages, 5 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 幻觉与事实性 5 篇

2505.16146 2025-09-16 cs.CV cs.AI cs.CL cs.LG 67%

Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation

Zhenglin Hua, Jinghan He, Zijun Yao, Tianxu Han, Haiyun Guo, Yuheng Jia, Junfeng Fang

机构 * School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University)(东南大学新一代人工智能技术及其交叉应用关键实验室) Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Wuhan University of Technology(武汉理工大学) National University of Singapore(新加坡国立大学)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15133 2025-09-16 cs.CL cs.AI cs.CV cs.HC cs.LG 67%

EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models

Ziwen Xu, Shuxun Wang, Kewei Xu, Haoming Xu, Mengru Wang, Xinle Deng, Yunzhi Yao, Guozhou Zheng, Huajun Chen, Ningyu Zhang

机构 * Zhejiang University(浙江大学) Ocean Research Center of Zhoushan, Zhejiang University(舟山海洋研究中心,浙江大学)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP 2025 System Demonstrations. Demo: https://www.youtube.com/watch?v=AkfoiPfp5rQ; code: https://github.com/zjunlp/EasyEdit

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07082 2025-09-16 cs.CV cs.AI cs.LG 62%

On the Generalization of Representation Uncertainty in Earth Observation

Spyros Kondylatos, Nikolaos Ioannis Bountos, Dimitrios Michail, Xiao Xiang Zhu, Gustau Camps-Valls, Ioannis Papoutsis

机构 * National Observatory of Athens(雅典国家天文台) National Technical University of Athens(雅典技术大学) University of Valencia(瓦伦西亚大学) Harokopio University of Athens(雅典惠克罗波利斯大学) Technical University of Munich(慕尼黑技术大学) Munich Center for Machine Learning(慕尼黑机器学习中心) Archimedes/Athena RC(阿基米德/雅典RC)

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted to ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01222 2025-09-16 cs.LG cs.AI 62%

Calibration in Deep Learning: A Survey of the State-of-the-Art

Cheng Wang

机构 * Amazon(亚马逊)

专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12034 2025-09-16 cs.AI 57%

Human-AI Use Patterns for Decision-Making in Disaster Scenarios: A Systematic Review

Emmanuel Adjei Domfeh, Christopher L. Dancy

机构 * Department of Computer Science and Engineering(计算机科学与工程系) The Pennsylvania State University(宾夕法尼亚州立大学) Department of Industrial and Manufacturing Engineering(工业与制造工程系)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 隐私与版权 1 篇

2509.10538 2025-09-16 cs.LG cs.AI cs.CL cs.CY 70%

DualAlign: Generating Clinically Grounded Synthetic Data

Rumeng Li, Xun Wang, Hong Yu

专题命中 隐私与版权 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 安全评测 17 篇

2509.11507 2025-09-16 cs.AI 70%

MedicalOS: An LLM Agent based Operating System for Digital Healthcare

Jared Zhu, Junde Wu

机构 * University of Oxford(牛津大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09686 2025-09-16 cs.IR cs.AI 70%

GeoGPT-RAG Technical Report

Fei Huang, Fan Wu, Zeqing Zhang, Qihao Wang, Long Zhang, Grant Michael Boquet, Hongyang Chen

机构 * GeoGPT Team(GeoGPT团队) Zhejiang Lab(浙江实验室)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI

Comments 19 pages, 10 figures, 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11595 2025-09-16 cs.AI cs.CE cs.CR cs.LG cs.MA 62%

AMLNet: A Knowledge-Based Multi-Agent Framework to Generate and Detect Realistic Money Laundering Transactions

Sabin Huda, Ernest Foo, Zahra Jadidi, MA Hakim Newton, Abdul Sattar

机构 * School of Information and Communication Technology, Griffith University, QLD Australia(信息与通信技术学院,格里菲斯大学,昆士兰州澳大利亚) School of Information and Physical Sciences, The University of Newcastle, NSW Australia(信息与物理科学学院,新castle大学,新南威尔士州澳大利亚)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏