arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9380 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9380 篇

2504.07995 2025-04-16 cs.CL 83%

SafeChat: A Framework for Building Trustworthy Collaborative Assistants and a Case Study of its Usefulness

Biplav Srivastava, Kausik Lakkaraju, Nitin Gupta, Vansh Nagpal, Bharath C. Muppasani, Sara E. Jones

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07424 2025-04-11 cs.AI 83%

Routing to the Right Expertise: A Trustworthy Judge for Instruction-based Image Editing

Chenxi Sun, Hongzhi Zhang, Qi Wang, Fuzheng Zhang

专题命中 安全评测 :trustworthy(title,abstract);alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.00393 2025-04-01 cs.AI 83%

Augmenting deep neural networks with symbolic knowledge: Towards trustworthy and interpretable AI for education

Danial Hooshyar, Roger Azevedo, Yeongwook Yang

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05475 2025-02-11 cs.LG 83%

You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation

Simon Pepin Lehalleur, Jesse Hoogland, Matthew Farrugia-Roberts, Susan Wei, Alexander Gietelink Oldenziel, George Wang, Liam Carroll, Daniel Murfet

专题命中 安全评测 :alignment(title,abstract);safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08088 2024-11-14 cs.CY cs.CR 83%

Safety case template for frontier AI: A cyber inability argument

Arthur Goemans, Marie Davidsen Buhl, Jonas Schuett, Tomek Korbak, Jessica Wang, Benjamin Hilton, Geoffrey Irving

专题命中 安全评测 :safety(title,abstract);AI safety(abstract);分类 cs.CY

Comments 21 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06835 2024-11-12 cs.CL cs.CR 83%

HarmLevelBench: Evaluating Harm-Level Compliance and the Impact of Quantization on Model Alignment

Yannis Belkhiter, Giulio Zizzo, Sergio Maffeis

专题命中 安全评测 :alignment(title,abstract);safety(abstract);分类 cs.CL

Comments NeurIPS 2024 Workshop on Safe Generative Artificial Intelligence (SafeGenAI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00069 2024-11-04 cs.CR cs.AI 83%

Meta-Sealing: A Revolutionizing Integrity Assurance Protocol for Transparent, Tamper-Proof, and Trustworthy AI System

Mahesh Vaijainthymala Krishnamoorthy

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI

Comments 24 pages, 3 figures and 10 Code blocks, to be presented in the conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.20087 2024-11-01 cs.LG cs.AI cs.CL cs.CY cs.HC 83%

ProgressGym: Alignment with a Millennium of Moral Progress

Tianyi Qiu, Yang Zhang, Xuchuan Huang, Jasmine Xinze Li, Jiaming Ji, Yaodong Yang

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments NeurIPS 2024 Track on Datasets and Benchmarks (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.04418 2024-10-28 cs.LG cs.CR cs.NI eess.SP 83%

Trustworthy Federated Learning via Blockchain

Zhanpeng Yang, Yuanming Shi, Yong Zhou, Zixin Wang, Kai Yang

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.LG

Comments This work has been submitted to the IEEE Internet of Things Journal for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19521 2024-10-01 cs.CR cs.LG 83%

GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks

Rongchang Li, Minjie Chen, Chang Hu, Han Chen, Wenpeng Xing, Meng Han

专题命中 安全评测 :prompt injection(title,abstract);safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04313 2024-07-15 cs.LG cs.AI cs.CL cs.CV cs.CY 83%

Improving Alignment and Robustness with Circuit Breakers

Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, Dan Hendrycks

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments Code and models are available at https://github.com/GraySwanAI/circuit-breakers

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03575 2024-06-21 cs.AI cs.HC 83%

Toward Human-AI Alignment in Large-Scale Multi-Player Games

Sugandha Sharma, Guy Davidson, Khimya Khetarpal, Anssi Kanervisto, Udit Arora, Katja Hofmann, Ida Momennejad

专题命中 安全评测 :alignment(title,abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19524 2024-05-31 cs.CR cs.AI 83%

AI Risk Management Should Incorporate Both Safety and Security

Xiangyu Qi, Yangsibo Huang, Yi Zeng, Edoardo Debenedetti, Jonas Geiping, Luxi He, Kaixuan Huang, Udari Madhushani, Vikash Sehwag, Weijia Shi, Boyi Wei, Tinghao Xie, Danqi Chen, Pin-Yu Chen, Jeffrey Ding, Ruoxi Jia, Jiaqi Ma, Arvind Narayanan, Weijie J Su, Mengdi Wang, Chaowei Xiao, Bo Li, Dawn Song, Peter Henderson, Prateek Mittal

专题命中 安全评测 :safety(title,abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14580 2024-02-23 cs.AI cs.SY eess.SY 83%

Savvy: Trustworthy Autonomous Vehicles Architecture

Ali Shoker, Rehana Yasmin, Paulo Esteves-Verissimo

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05818 2023-10-10 cs.CL 83%

SC-Safety: A Multi-round Open-ended Question Adversarial Safety Benchmark for Large Language Models in Chinese

Liang Xu, Kangkang Zhao, Lei Zhu, Hang Xue

专题命中 安全评测 :safety(title,abstract);trustworthy(abstract);分类 cs.CL

Comments 20 pages, 8 tables, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.06421 2023-10-06 cs.AI cs.HC 83%

AI Alignment Dialogues: An Interactive Approach to AI Alignment in Support Agents

Pei-Yu Chen, Myrthe L. Tielman, Dirk K. J. Heylen, Catholijn M. Jonker, M. Birna van Riemsdijk

专题命中 安全评测 :alignment(title,abstract);trustworthy(abstract);分类 cs.AI

Comments Withdraw because the content of the paper has been largely revised. The newest version is very different than the submitted one

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.16144 2023-08-09 cs.RO cs.AI 83%

Towards trustworthy multi-modal motion prediction: Holistic evaluation and interpretability of outputs

Sandra Carrasco Limeros, Sylwia Majchrowska, Joakim Johnander, Christoffer Petersson, Miguel Ángel Sotelo, David Fernández Llorca

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI

Comments 16 pages, 7 figures, 6 tables

Journal ref CAAI Transactions on Intelligence Technology 24686557 (ISSN) 24682322 (eISSN) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09705 2023-07-20 cs.CL 83%

CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility

Guohai Xu, Jiayi Liu, Ming Yan, Haotian Xu, Jinghui Si, Zhuoran Zhou, Peng Yi, Xing Gao, Jitao Sang, Rong Zhang, Ji Zhang, Chao Peng, Fei Huang, Jingren Zhou

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CL

Comments Working in Process

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.10007 2023-03-20 cs.CY 83%

Information Governance as a Socio-Technical Process in the Development of Trustworthy Healthcare AI

Nigel Rees, Kelly Holding, Mark Sujan

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.CY

Journal ref Frontiers in Computer Science. 2023;5

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.14451 2022-09-13 cs.CY 83%

Foreseeing the Impact of the Proposed AI Act on the Sustainability and Safety of Critical Infrastructures

Francesco Sovrano, Giulio Masetti

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CY

Comments 10 pages, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.08580 2022-01-24 cs.IR cs.AI cs.DB 83%

Trustworthy Knowledge Graph Completion Based on Multi-sourced Noisy Data

Jiacheng Huang, Yao Zhao, Wei Hu, Zhen Ning, Qijin Chen, Xiaoxia Qiu, Chengfu Huo, Weijun Ren

专题命中 安全评测 :trustworthy(title,abstract);alignment(abstract);分类 cs.AI

Comments Accepted in the ACM Web Conference (WWW 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.14019 2021-10-28 cs.LG 83%

Reliable and Trustworthy Machine Learning for Health Using Dataset Shift Detection

Chunjong Park, Anas Awadalla, Tadayoshi Kohno, Shwetak Patel

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.LG

Comments Neu

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.06641 2021-08-20 cs.AI 83%

Trustworthy AI: A Computational Perspective

Haochen Liu, Yiqi Wang, Wenqi Fan, Xiaorui Liu, Yaxin Li, Shaili Jain, Yunhao Liu, Anil K. Jain, Jiliang Tang

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI

Comments 55 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.04408 2021-05-11 cs.RO cs.AI 83%

The Challenges and Opportunities of Human-Centered AI for Trustworthy Robots and Autonomous Systems

Hongmei He, John Gray, Angelo Cangelosi, Qinggang Meng, T. Martin McGinnity, Jörn Mehnen

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.05501 2020-09-14 stat.ML cs.LG 83%

Towards a More Reliable Interpretation of Machine Learning Outputs for Safety-Critical Systems using Feature Importance Fusion

Divish Rengasamy, Benjamin Rothwell, Grazziela Figueredo

专题命中 安全评测 :safety(title,abstract);trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.01540 2019-04-03 cs.AI 83%

Augmented Utilitarianism for AGI Safety

Nadisha-Marie Aliman, Leon Kester

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07744 2024-02-15 cs.AI cs.CL cs.LG 83%

Towards Unified Alignment Between Agents, Humans, and Environment

Zonghan Yang, An Liu, Zijun Liu, Kaiming Liu, Fangzhou Xiong, Yile Wang, Zeyuan Yang, Qingyuan Hu, Xinrui Chen, Zhenhe Zhang, Fuwen Luo, Zhicheng Guo, Peng Li, Yang Liu

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Project webpage: https://agent-force.github.io/unified-alignment-for-agents.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12382 2026-04-28 cs.LG cs.AI cs.CR 82%

Exploring the Secondary Risks of Large Language Models

探索大型语言模型的二次风险

Jiawei Chen, Zhengwei Fang, Yu Tian, Jiawei Du, Chao Yu, Zhaoxia Yin, Hang Su

机构 * Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院,清华大学) Beijing Zhongguancun Academy(北京中关村学院) Department of Computer Science and Technology, THBI Lab, Tsinghua University(计算机科学与技术系,清华大学THBI实验室) CFAR, A*STAR, Singapore(新加坡A*STAR CFAR)

专题命中 安全评测 :jailbreak(abstract,abstract_cn);alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出二次风险作为新型失败模式,通过SecLens框架系统评估,揭示了大型语言模型在良性交互中存在广泛且转移性强的有害行为,强调需加强安全机制。

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01317 2026-08-25 cs.SE cs.CR 版本更新 82%

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces

SABER:在有状态项目工作区中对LLM编码代理的操作安全性进行基准测试

Qi Hu, Yifeng Tang, Qinghua Wang, Lanyang Zhao, Pengji Zhang, Yuhao Qing, Xin Yao, Dong Huang, Lin Zhang, Zhuoran Ji

专题命中 安全评测 :safety(title,abstract);alignment(abstract)

AI总结 提出SABER基准,通过将模型置于真实代理式项目中并评估最终环境状态,来衡量LLM编码代理的操作安全性,发现最佳模型的有害安全违规率超过54%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04806 2026-08-25 cs.LG cs.AI cs.CL 版本更新 82%

NeST: Neighborhood-aware semantic alignment and temporal modulation for LLM based time series forecasting

NeST:面向基于大语言模型的时间序列预测的邻域感知语义对齐与时间调制

Jayanie Bogahawatte, Sachith Seneviratne, Maneesha Perera, Saman Halgamuge

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 NeST框架通过邻域感知语义对齐与时间调制构建文本整合提示词微调LLM,在8个基准及9个真实光伏数据集的时间序列预测中,MSE、R²等指标优于现有方法,泛化性强。

详情

展开后加载摘要…

URL PDF HTML 收藏