arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-12 至 2025-08-12 共收录 77 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 7 篇

2508.07768 2025-08-12 cs.LG cs.AI cs.CL 85%

Pareto Multi-Objective Alignment for Language Models

Qiang He, Setareh Maghsudi

机构 * Ruhr University Bochum(鲁尔大学波恩)

专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at ECML/PKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07750 2025-08-12 cs.LG cs.AI cs.CL 85%

Learning to Align, Aligning to Learn: A Unified Approach for Self-Optimized Alignment

Haowen Wang, Yun Yue, Zhiling Ye, Shuowen Zhang, Lei Fan, Jiaxin Liang, Jiadi Jiang, Cheng Wei, Jingyuan Deng, Xudong Han, Ji Li, Chunxiao Guo, Peng Wei, Jian Wang, Jinjie Gu

专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 12 pages, 5 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23998 2025-08-12 cs.CL 70%

Auto-TA: Towards Scalable Automated Thematic Analysis (TA) via Multi-Agent Large Language Models with Reinforcement Learning

Seungjun Yi, Joakim Nguyen, Huimin Xu, Terence Lim, Andrew Well, Mia Markey, Ying Ding

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract);分类 cs.CL

Comments Presented at ACL 2025 SRW

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00624 2025-08-12 cs.CV 67%

VideoSAVi: Self-Aligned Video Language Models without Human Supervision

Yogesh Kulkarni, Pooyan Fazli

机构 * Arizona State University(亚利桑那州立大学)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract)

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06963 2025-08-12 cs.AI cs.LG 62%

MASteer: Multi-Agent Adaptive Steer Strategy for End-to-End LLM Trustworthiness Repair

Changqing Li, Tianlin Li, Xiaohan Zhang, Aishan Liu, Li Pan

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08101 2025-08-12 cs.HC cs.AI cs.SE 57%

ChatGPT on the Road: Leveraging Large Language Model-Powered In-vehicle Conversational Agents for Safer and More Enjoyable Driving Experience

Yeana Lee Bond, Mungyeong Choe, Baker Kasim Hasan, Arsh Siddiqui, Myounghoon Jeon

机构 * Computer Science, Virginia Tech, Blacksburg, Virginia, USA(计算机科学,弗吉尼亚理工学院,弗吉尼亚州黑斯堡,美国;弗吉尼亚理工学院黑斯堡弗吉尼亚美国) Virginia Tech Blacksburg Virginia USA

专题命中 偏好对齐 :safety(abstract);分类 cs.AI

Comments Submitted to International Journal of Human-Computer Studies. Bond and Choe: Drafting, Review, Editing, Validation, Software, Methodology, Investigation, Data Analysis, Conceptualization, Experiment training. Hasan and Siddiqui: Experimental and Data Analysis Support. Jeon: Supervision, Review, Resources, Project Admin, Methodology, Conceptualization. Total 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14860 2025-08-12 cs.CL 57%

ALFA: Aligning LLMs to Ask Good Questions A Case Study in Clinical Reasoning

Shuyue Stella Li, Jimin Mun, Faeze Brahman, Pedram Hosseini, Bryceton G. Thomas, Jessica M. Sin, Bing Ren, Jonathan S. Ilgen, Yulia Tsvetkov, Maarten Sap

机构 * University of Washington(华盛顿大学) Carnegie Mellon University(卡内基梅隆大学) Allen Institute for AI(人工智能研究院) Lavita AI(Lavita人工智能) Dartmouth Medicine(达特茅斯医学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments 29 pages, 8 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全训练 12 篇

2412.16633 2025-08-12 cs.RO cs.AI cs.CY 82%

POEX: Towards Policy Executable Jailbreak Attacks Against the LLM-based Robots

Xuancun Lu, Zhengxian Huang, Xinfeng Li, Chi Zhang, Xiaoyu ji, Wenyuan Xu

专题命中 安全训练 :jailbreak(title,abstract);分类 cs.AI、cs.CY

Comments Homepage: https://poex-jailbreak.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13774 2025-08-12 cs.AI cs.CY cs.MA 73%

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values

Nell Watson, Ahmed Amer, Evan Harris, Preeti Ravindra, Shujun Zhang

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY

Comments 42 pages, 6 figures

Journal ref Information 2025, 16(8), 651

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07651 2025-08-12 eess.SP 71%

Remote ID Based UAV Collision Avoidance Optimization for Low-Altitude Airspace Safety

Ziye Jia, Yian Zhu, Qihui Wu, Lei Zhang, Sen Yang, Zhu Han

专题命中 安全训练 :safety(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07172 2025-08-12 cs.CL 70%

Gradient Surgery for Safe LLM Fine-Tuning

Biao Yi, Jiahao Li, Baolei Zhang, Lihai Nie, Tong Li, Tiansheng Huang, Zheli Liu

机构 * Nankai University(南开大学)

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17130 2025-08-12 cs.CL cs.CR cs.CY 62%

Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control

Hannah Cyberey, David Evans

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.CY

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07075 2025-08-12 cs.LG cs.AI 62%

Surgical Knowledge Rewrite in Compact LLMs: An 'Unlearn-then-Learn' Strategy with ($IA^3$) for Localized Factual Modulation and Catastrophic Forgetting Mitigation

Stanley Ngugi

机构 * Stanley Ngugi(独立研究者)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 9 pages, 2 visual aids

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19250 2025-08-12 cs.LG cs.AI 62%

Robust Behavior Cloning Via Global Lipschitz Regularization

Shili Wu, Yizhao Jin, Puhua Niu, Aniruddha Datta, Sean B. Andersson

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07096 2025-08-12 cs.AI 57%

Rejecting Hallucinated State Targets during Planning

Mingde Zhao, Tristan Sylvain, Romain Laroche, Doina Precup, Yoshua Bengio

机构 * McGill University(麦吉尔大学) Mila (Quebec AI Institute)(蒙特利尔AI研究所) Google Deepmind(谷歌DeepMind)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments [20250810]: ICML 2025 Camera Ready, https://github.com/mila-iqia/delusions

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06754 2025-08-12 cs.AI 57%

A Fuzzy Logic Prompting Framework for Large Language Models in Adaptive and Uncertain Tasks

Vanessa Figueiredo

专题命中 安全训练 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00352 2025-08-12 cs.AI cs.MA cs.RO 57%

A Differentiated Reward Method for Reinforcement Learning based Multi-Vehicle Cooperative Decision-Making Algorithms

Ye Han, Lijun Zhang, Dejian Meng, Zhuang Zhang

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22792 2025-08-12 cs.CV 50%

Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization

Yuxi Zhang, Yueting Li, Xinyu Du, Sibo Wang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of California, Berkeley(加州大学伯克利分校)

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.24152 2025-08-12 cs.RO 50%

Language-Driven Policy Distillation for Cooperative Driving in Multi-Agent Reinforcement Learning

Jiaqi Liu, Chengkai Xu, Peng Hang, Jian Sun, Wei Zhan, Masayoshi Tomizuka, Mingyu Ding

机构 * Department of Computer Science at University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校计算机科学系) College of Transportation and Key Laboratory of Road and Traffic Engineering, Ministry of Education, Tongji University(同济大学交通学院及交通工程教育部重点实验室) Department of Mechanical Engineering at the University of California, Berkeley(加州大学伯克利分校机械工程系)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 越狱攻击 4 篇

2505.17066 2025-08-12 cs.CR cs.AI 83%

Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration

Tatia Tsmindashvili, Ana Kolkhidashvili, Dachi Kurtskhalia, Nino Maghlakelidze, Elene Mekvabishvili, Guram Dentoshvili, Orkhan Shamilov, Zaal Gachechiladze, Steven Saporta, David Dachi Choladze

专题命中 越狱攻击 :jailbreak(title,abstract);prompt injection(abstract);分类 cs.AI

Journal ref IEEE Access, vol. 13, pp. 134976-134988, Jul. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07646 2025-08-12 cs.LG 77%

Multi-Turn Jailbreaks Are Simpler Than They Seem

Xiaoxue Yang, Jaeha Lee, Anna-Katharina Dick, Jasper Timm, Fei Xie, Diogo Cruz

机构 * Imperial College London(伦敦帝国学院) California Institute of Technology(加州理工学院)

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);AI safety(abstract);分类 cs.LG

Comments 25 pages, 15 figures. Accepted at COLM 2025 SoLaR Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06755 2025-08-12 cs.CL cs.AI 73%

Many-Turn Jailbreaking

Xianjun Yang, Liqiang Xiao, Shiyang Li, Faisal Ladhak, Hyokun Yun, Linda Ruth Petzold, Yi Xu, William Yang Wang

专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07139 2025-08-12 cs.CR cs.AI 70%

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection

Ivan Zhang

机构 * CMU(卡内基梅隆大学)

专题命中 越狱攻击 :alignment(abstract);jailbreak(abstract);分类 cs.AI

Comments 10 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 1 篇

2508.07223 2025-08-12 cs.IR cs.AI 57%

Selection and Exploitation of High-Quality Knowledge from Large Language Models for Recommendation

Guanchen Wang, Mingming Ha, Tianbao Ma, Linxun Chen, Zhaojie Liu, Guorui Zhou, Kun Gai

专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 2 篇

2508.07672 2025-08-12 cs.HC 50%

Towards Aligning Personalized Conversational Recommendation Agents with Users' Privacy Preferences

Shuning Zhang, Ying Ma, Jingruo Chen, Simin Li, Xin Yi, Hewu Li

专题命中 隐私与版权 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07664 2025-08-12 cs.HC 50%

Understanding Users' Privacy Perceptions Towards LLM's RAG-based Memory

Shuning Zhang, Rongjun Ma, Ying Ma, Shixuan Li, Yiqun Xu, Xin Yi, Hewu Li

专题命中 隐私与版权 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 16 篇

2508.07390 2025-08-12 cs.HC cs.AI 79%

Urbanite: A Dataflow-Based Framework for Human-AI Interactive Alignment in Urban Visual Analytics

Gustavo Moreira, Leonardo Ferreira, Carolina Veiga, Maryam Hosseini, Fabio Miranda

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments Accepted at IEEE VIS 2025. Urbanite is available at https://urbantk.org/urbanite

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07031 2025-08-12 eess.IV cs.AI cs.CV 79%

Trustworthy Medical Imaging with Large Language Models: A Study of Hallucinations Across Modalities

Anindya Bijoy Das, Shahnewaz Karim Sakib, Shibbir Ahmed

机构 * The University of Akron(阿克隆大学) University of Tennessee at Chattanooga(田纳西大学查塔努加分校) Texas State University(德克萨斯州立大学)

专题命中 安全评测 :trustworthy(title);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06885 2025-08-12 cs.LG 79%

Conformal Prediction and Trustworthy AI

Anthony Bellotti, Xindi Zhao

机构 * School of Computer Science, University of Nottingham Ningbo China(诺丁汉大学宁波校区计算机科学学院)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

Comments Preprint for an essay to be published in The Importance of Being Learnable (Enhancing the Learnability and Reliability of Machine Learning Algorithms) Essays Dedicated to Alexander Gammerman on His 80th Birthday, LNCS Springer Nature Switzerland AG ed. Nguyen K.A. and Luo Z

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07560 2025-08-12 cs.RO cs.CV 78%

Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

Yan Gong, Naibang Wang, Jianli Lu, Xinyu Zhang, Yongsheng Gao, Jie Zhao, Zifan Huang, Haozhi Bai, Nanxin Zeng, Nayu Su, Lei Yang, Ziying Song, Xiaoxi Hu, Xinmin Jiang, Xiaojuan Zhang, Susanto Rahardja

机构 * State Key Laboratory of Robotics and System(机器人系统国家重点实验室) Harbin Institute of Technology(哈尔滨工业大学) State Key Laboratory of Intelligent Green Vehicle and Mobility(智能绿色车辆与移动性国家重点实验室) Tsinghua University(清华大学) the School of Mechanical and Aerospace Engineering(机械与航空航天工程学院) Nanyang Technological University(南洋理工大学) Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence(北京交通数据挖掘与具身智能重点实验室) Beijing Jiaotong University(北京交通大学) the Institute for Infocomm Research(信息通信研究所) A*STAR the Engineering Cluster(工程集群) the Singapore Institute of Technology(新加坡理工学院)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏