arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3309 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3309 篇

2501.17070 2025-04-18 cs.CR cs.CL cs.LG 62%

Contextual Agent Security: A Policy for Every Purpose

Lillian Tsai, Eugene Bagdasarian

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.LG

Comments Workshop in Hot Topics in Operating Systems (HotOS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12299 2025-04-17 cs.AI cs.CV cs.LG 62%

Adapting a World Model for Trajectory Following in a 3D Game

Marko Tot, Shu Ishida, Abdelhak Lemkhenter, David Bignell, Pallavi Choudhury, Chris Lovett, Luis França, Matheus Ribeiro Furtado de Mendonça, Tarun Gupta, Darren Gehring, Sam Devlin, Sergio Valcarcel Macua, Raluca Georgescu

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12082 2025-04-17 cs.CL cs.AI 62%

Selective Demonstration Retrieval for Improved Implicit Hate Speech Detection

Yumin Kim, Hwanhee Lee

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12808 2025-04-11 cs.AI cs.CY cs.HC 62%

Conversational Medical AI: Ready for Practice

Antoine Lizée, Pierre-Auguste Beaucoté, James Whitbeck, Marion Doumeingts, Anaël Beaugnon, Isabelle Feldhaus

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.CY

Comments Accepted to AAAI25 (Oral, workshop) 14 pages, 7 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04215 2025-04-08 cs.CL cs.AI 62%

Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability

Vishnu Kabir Chhabra, Mohammad Mahdi Khalili

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21504 2025-03-28 cs.CL cs.AI cs.CV 62%

Keyword-Oriented Multimodal Modeling for Euphemism Identification

Yuxue Hu, Junsong Li, Meixuan Chen, Dongyu Su, Tongguan Wang, Ying Sha

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07671 2025-03-26 stat.ML cs.AI cs.LG 62%

Probabilistic Shielding for Safe Reinforcement Learning

Edwin Hamel-De le Court, Francesco Belardinelli, Alexander W. Goodall

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 13 pages, 3 figures, Conference: AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.12198 2025-03-25 cs.CL cs.CV cs.LG 62%

Exploring the Limits of Zero Shot Vision Language Models for Hate Meme Detection: The Vulnerabilities and their Interpretations

Naquee Rizwan, Paramananda Bhaskar, Mithun Das, Swadhin Satyaprakash Majhi, Punyajoy Saha, Animesh Mukherjee

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.20089 2025-03-21 cs.LG cs.CL cs.CR 62%

Robust LLM safeguarding via refusal feature adversarial training

Lei Yu, Virginie Do, Karen Hambardzumyan, Nicola Cancedda

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14703 2025-03-20 cs.CL cs.AI 62%

Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics

Seungbeen Lee, Seungwon Lim, Seungju Han, Giyeong Oh, Hyungjoo Chae, Jiwan Chung, Minju Kim, Beong-woo Kwak, Yeonsoo Lee, Dongha Lee, Jinyoung Yeo, Youngjae Yu

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to NAACL2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12753 2025-03-18 cs.NI cs.AI cs.LG 62%

SafeSlice: Enabling SLA-Compliant O-RAN Slicing via Safe Deep Reinforcement Learning

Ahmad M. Nagib, Hatem Abou-Zeid, Hossam S. Hassanein

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments This article has been accepted for presentation in the IEEE International Conference on Machine Learning for Communication and Networking (ICMLCN) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11696 2025-03-18 eess.SY cs.AI cs.LG cs.SY 62%

Balancing SoC in Battery Cells using Safe Action Perturbations

E Harshith Kumar Yadav, Rahul Narava, Anshika, Shashi Shekher Jha

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13187 2025-03-11 cs.LG cs.AI cs.RO 62%

A Survey of Sim-to-Real Methods in RL: Progress, Prospects and Challenges with Foundation Models

Longchao Da, Justin Turnau, Thirulogasankar Pranav Kutralingam, Alvaro Velasquez, Paulo Shakarian, Hua Wei

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 19 pages, 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10998 2025-03-06 eess.SY cs.AI cs.LG cs.LO cs.SY 62%

Provably Safe Neural Network Controllers via Differential Dynamic Logic

Samuel Teuber, Stefan Mitsch, André Platzer

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 39 pages (main paper has 10 pages), 13 figures; Accepted at the Thirty-Eighth Annual Conference on Neural Information Processing Systems (NeurIPS 2024)

Journal ref in Advances in Neural Information Processing Systems, 2024, pp. 1586-1624

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.19271 2025-02-28 cs.CL cs.AI cs.IR 62%

AutoPureData: Automated Filtering of Undesirable Web Data to Update LLM Knowledge

Praneeth Vadlapati

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments Final version

Journal ref Journal of Mathematical & Computer Applications, 3 (2024) E121

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12411 2025-02-19 cs.CL cs.AI 62%

Gradient Co-occurrence Analysis for Detecting Unsafe Prompts in Large Language Models

Jingyuan Yang, Bowen Yan, Rongjun Li, Ziyu Zhou, Xin Chen, Zhiyong Feng, Wei Peng

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14853 2025-02-18 cs.CL cs.AI cs.CR 62%

Atoxia: Red-teaming Large Language Models with Target Toxic Answers

Yuhao Du, Zhuo Li, Pengyu Cheng, Xiang Wan, Anningzhe Gao

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to Findings of NAACL-2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10431 2025-02-18 cs.LG cs.AI 62%

Leveraging Constraint Violation Signals For Action-Constrained Reinforcement Learning

Janaka Chathuranga Brahmanage, Jiajing Ling, Akshat Kumar

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 7 pages and 5 pages supplementary

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09886 2025-02-17 cs.RO cs.AI cs.LG 62%

Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos

Weirui Ye, Fangchen Liu, Zheng Ding, Yang Gao, Oleh Rybkin, Pieter Abbeel

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19471 2025-02-17 cs.RO cs.AI cs.CL cs.FL 62%

SELP: Generating Safe and Efficient Task Plans for Robot Agents with Large Language Models

Yi Wu, Zikang Xiong, Yiran Hu, Shreyash S. Iyengar, Nan Jiang, Aniket Bera, Lin Tan, Suresh Jagannathan

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

Comments This paper has been accepted for presentation at the 2025 IEEE International Conference on Robotics and Automation (ICRA), May 19-23, 2025, Atlanta, USA, and for inclusion in the conference proceeding

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04184 2025-02-17 cs.LO cs.AI cs.LG cs.RO 62%

Shield Synthesis for LTL Modulo Theories

Andoni Rodriguez, Guy Amir, Davide Corsi, Cesar Sanchez, Guy Katz

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments To appear in AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09175 2025-02-14 cs.CR cs.AI cs.CL 62%

FLAME: Flexible LLM-Assisted Moderation Engine

Ivan Bakulin, Ilia Kopanichuk, Iaroslav Bespalov, Nikita Radchenko, Vladimir Shaposhnikov, Dmitry Dylov, Ivan Oseledets

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04249 2025-02-07 cs.AI cs.LG cs.MA physics.data-an stat.ML 62%

Free Energy Risk Metrics for Systemically Safe AI: Gatekeeping Multi-Agent Study

Michael Walters, Rafael Kaufmann, Justice Sefas, Thomas Kopinski

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments 9 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17460 2025-02-07 cs.RO cs.AI cs.CV cs.LG cs.SY eess.SY 62%

SoNIC: Safe Social Navigation with Adaptive Conformal Inference and Constrained Reinforcement Learning

Jianpeng Yao, Xiaopan Zhang, Yu Xia, Zejin Wang, Amit K. Roy-Chowdhury, Jiachen Li

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments Project website: https://sonic-social-nav.github.io/; 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01088 2025-02-05 cs.HC cs.CL cs.LG 62%

Exploring Empty Spaces: Human-in-the-Loop Data Augmentation

Catherine Yeh, Donghao Ren, Yannick Assogba, Dominik Moritz, Fred Hohman

专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.LG

Comments Code: https://github.com/apple/ml-interactive-data-augmentation/

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15844 2025-02-03 stat.ML cs.AI cs.IT cs.LG math.IT stat.ME 62%

Adaptive Learn-then-Test: Statistically Valid and Efficient Hyperparameter Selection

Matteo Zecchin, Sangwoo Park, Osvaldo Simeone

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17217 2025-01-28 cs.LG cs.AI 62%

Temporal Logic Specification-Conditioned Decision Transformer for Offline Safe Reinforcement Learning

Zijian Guo, Weichao Zhou, Wenchao Li

专题命中 安全训练 :alignment(abstract);分类 cs.AI、cs.LG

Comments Published in ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08145 2025-01-15 cs.CL cs.AI 62%

Refusal Behavior in Large Language Models: A Nonlinear Perspective

Fabian Hildebrandt, Andreas Maier, Patrick Krauss, Achim Schilling

专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.05113 2025-01-10 physics.comp-ph cs.AI cs.LG 62%

Constrained Optimization of Charged Particle Tracking with Multi-Agent Reinforcement Learning

Tobias Kortus, Ralf Keidel, Nicolas R. Gauger, Jan Kieseler

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.04982 2025-01-10 cs.RO cs.AI cs.LG 62%

CuRLA: Curriculum Learning Based Deep Reinforcement Learning for Autonomous Driving

Bhargava Uppuluri, Anjel Patel, Neil Mehta, Sridhar Kamath, Pratyush Chakraborty

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.LG

Comments To be published in the 17th International Conference on Agents and Artificial Intelligence (ICAART), Feb 2025

详情

展开后加载摘要…

URL PDF HTML 收藏