arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9400 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9400 篇

2406.12624 2025-08-19 cs.CL cs.AI 81%

Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, Dieuwke Hupkes

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Meta

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments https://aclanthology.org/2025.gem-1.33/

Journal ref Proceedings of the Fourth Workshop on Generation Evaluation and Metrics GEM2 2025 pages 404 to 430; July 31 August 1 2025; 2025 Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13983 2025-08-13 cs.CL cs.AI 81%

AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models

Yang Fan

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments There are serious academic problems in this paper, such as data falsification and plagiarism in the method of the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02078 2025-08-13 cs.CV cs.AI cs.LG 81%

From Lab to Field: Real-World Evaluation of an AI-Driven Smart Video Solution to Enhance Community Safety

Shanle Yao, Babak Rahimi Ardabili, Armin Danesh Pazho, Ghazal Alinezhad Noghre, Christopher Neff, Lauren Bourque, Hamed Tabkhi

机构 * Department of Electrical and Computer Engineering, University of North Carolina at Charlotte(电气与计算机工程系,北卡罗来纳大学夏洛特分校) Department of Public Policy, University of North Carolina at Charlotte(公共政策系,北卡罗来纳大学夏洛特分校)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08273 2025-08-13 cs.CL cs.LG 81%

TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning

Kristian Miok, Blaz Škrlj, Daniela Zaharie, Marko Robnik Šikonja

机构 * Faculty of Computer and Information Science, University of Ljubljana, Slovenia(卢布尔雅那大学计算机与信息科学学院) ICAM - Advanced Environmental Research Institute, West University of Timisoara, Romania(蒂米șoara西大学先进环境研究所) Department of Computer Science, West University of Timisoara, Romania(蒂米șoa拉西大学计算机科学系)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16591 2025-07-31 cs.LG cs.AI cs.CR 81%

Bridging Privacy and Robustness for Trustworthy Machine Learning

Xiaojin Zhang, Wei Chen

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17616 2025-07-24 cs.CV cs.AI cs.LG 81%

Vision Transformer attention alignment with human visual perception in aesthetic object evaluation

Miguel Carrasco, César González-Martín, José Aranda, Luis Oliveros

机构 * Escuela de Informática y Telecomunicaciones, Universidad Diego Portáles(戴维·波特莱斯大学信息与电信学院) Department of Specific Didactics, University of Cordoba(科尔多瓦大学特定教学系) Facultad de Ingeniería y Ciencias, Universidad Adolfo Ibáñez(阿道弗·伊巴涅斯大学工程与科学学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 25 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12404 2025-07-24 cs.LG cs.AI 81%

EXGnet: a single-lead explainable-AI guided multiresolution network with train-only quantitative features for trustworthy ECG arrhythmia classification

Tushar Talukder Showrav, Soyabul Islam Lincoln, Md. Kamrul Hasan

机构 * Dept. of Electrical and Electronic Engineering(电子与电气工程系) Bangladesh University of Engineering and Technology(孟加拉工程与技术大学) Dept. of Electronics & Communication Engineering(电子与通信工程系) Khulna University of Engineering and Technology(库尔纳工程与技术大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02145 2025-07-21 cs.AI cs.CL cs.RO 81%

From Words to Collisions: LLM-Guided Evaluation and Adversarial Generation of Safety-Critical Driving Scenarios

Yuan Gao, Mattia Piccinini, Korbinian Moller, Amr Alanwar, Johannes Betz

机构 * Professorship of Autonomous Vehicle Systems, TUM School of Engineering and Design, Technical University of Munich(自主车辆系统教授职位,技术大学慕尼黑工程与设计学院) Munich Institute of Robotics and Machine Intelligence (MIRMI)(慕尼黑机器人与机器智能研究所) TUM School of Computation, Information and Technology, Department of Computer Engineering, Technical University of Munich(技术大学慕尼黑计算、信息与技术学院,计算机工程系)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

Comments Final Version and Paper Accepted at IEEE ITSC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17352 2025-07-03 cs.CL cs.AI 81%

Towards Safety Evaluations of Theory of Mind in Large Language Models

Tatsuhiro Aoshima, Mitsuaki Akiyama

机构 * NTT(日本电报电话株式会社)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19564 2025-07-01 cs.LG cs.AI 81%

FedMM-X: A Trustworthy and Interpretable Framework for Federated Multi-Modal Learning in Dynamic Environments

Sree Bhargavi Balija

机构 * University of California, San Diego(加州大学圣地亚哥分校)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22446 2025-07-01 cs.LG cs.AI 81%

EAGLE: Efficient Alignment of Generalized Latent Embeddings for Multimodal Survival Prediction with Interpretable Attribution Analysis

Aakash Tripathi, Asim Waqas, Matthew B. Schabath, Yasin Yilmaz, Ghulam Rasool

机构 * Dept. of Machine Learning Moffitt Cancer Center(机器学习系莫菲特癌症中心) Dept. of Cancer Epidemiology Moffitt Cancer Center(癌症流行病学系莫菲特癌症中心) Dept. of Electrical Engineering University of South Florida(电气工程系佛罗里达州立大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21252 2025-06-27 cs.CL cs.AI 81%

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents

Tianyi Men, Zhuoran Jin, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

Comments ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20251 2025-06-26 cs.LG cs.AI 81%

Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models

Kejia Chen, Jiawen Zhang, Jiacong Hu, Yu Wang, Jian Lou, Zunlei Feng, Mingli Song

机构 * The State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学) Sun Yat-sen University(中山大学) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术区(滨江)区块链与数据安全研究院)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14922 2025-06-26 cs.CY cs.LG 81%

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety

Christina Q. Knight, Kaustubh Deshpande, Ved Sirdeshmukh, Meher Mankikar, Scale Red Team, SEAL Research Team, Julian Michael

机构 * Scale AI

专题命中 安全评测 :safety(title,abstract);分类 cs.CY、cs.LG

Comments 12 pages, 7 figures, submitted to NeurIPS

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12368 2025-06-18 cs.CL cs.AI 81%

CAPTURE: Context-Aware Prompt Injection Testing and Robustness Enhancement

Gauri Kholkar, Ratinder Ahuja

机构 * First Author Affiliation(第一作者机构) Second Author Affiliation(第二作者机构)

专题命中 安全评测 :prompt injection(title,abstract);分类 cs.CL、cs.AI

Comments Accepted in ACL LLMSec Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10604 2025-06-18 cs.CV cs.AI cs.CL 81%

When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis

Ruixuan Zhang, Beichen Wang, Juexiao Zhang, Zilin Bian, Chen Feng, Kaan Ozbay

机构 * New York University(纽约大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04951 2025-06-16 cs.CR cs.AI cs.LG 81%

Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search

Zeren Luo, Zifan Peng, Yule Liu, Zhen Sun, Mingchen Li, Jingyi Zheng, Xinlei He

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07564 2025-06-12 cs.AI cs.CL 81%

SAFEFLOW: A Principled Protocol for Trustworthy and Transactional Autonomous Agent Systems

Peiran Li, Xinkai Zou, Zhuohang Wu, Ruifeng Li, Shuo Xing, Hanwen Zheng, Zhikai Hu, Yuping Wang, Haoxi Li, Qin Yuan, Yingmo Zhang, Zhengzhong Tu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

Comments Former versions either contain unrelated content or cannot be properly converted to PDF

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01080 2025-06-09 cs.AI cs.CY 81%

The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process

Florian Carichon, Aditi Khandelwal, Marylou Fauchard, Golnoosh Farnadi

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments Preprint of NeurIPS 2025 Position Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10640 2025-06-04 cs.SE cs.AI cs.LG 81%

The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware)

Kirill Vasilevski, Benjamin Rombaut, Gopi Krishnan Rajbahadur, Gustavo A. Oliva, Keheliya Gallaba, Filipe R. Cogo, Jiahuei Lin, Dayi Lin, Haoxiang Zhang, Bouyan Chen, Kishanthan Thangarajah, Ahmed E. Hassan, Zhen Ming Jiang

机构 * Centre for Software Excellence, Huawei Canada(华为加拿大软件卓越中心) Queen’s University, Canada(加拿大女王大学) York University, Canada(加拿大约克大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10927 2025-06-04 cs.CL cs.AI 81%

OASST-ETC Dataset: Alignment Signals from Eye-tracking Analysis of LLM Responses

Angela Lopez-Cardona, Sebastian Idesis, Miguel Barreda-Ángeles, Sergi Abadal, Ioannis Arapakis

机构 * Telefónica Scientific Research(Telefónica科学研究中心) Universitat Politècnica de Catalunya(巴塞罗那理工大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments This paper has been accepted to ACM ETRA 2025 and published on PACMHCI

Journal ref Proceedings of the ACM on Human-Computer Interaction. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05873 2025-06-03 cs.CL cs.AI 81%

MEXA: Multilingual Evaluation of English-Centric LLMs via Cross-Lingual Alignment

Amir Hossein Kargaran, Ali Modarressi, Nafiseh Nikeghbal, Jana Diesner, François Yvon, Hinrich Schütze

机构 * LMU Munich & Munich Center for Machine Learning(慕尼黑大学及慕尼黑机器学习中心) Technical University of Munich(慕尼黑技术大学) Sorbonne Université & CNRS, ISIR(索邦大学及CNRS,ISIR)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments ACL Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20144 2025-05-27 cs.CL cs.LG 81%

SeMe: Training-Free Language Model Merging via Semantic Alignment

Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu Zhang

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments an early-stage version

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.08085 2025-05-23 cs.CL cs.AI 81%

Can Knowledge Graphs Make Large Language Models More Trustworthy? An Empirical Study Over Open-ended Question Answering

Yuan Sui, Yufei He, Zifeng Ding, Bryan Hooi

机构 * National University of Singapore(新加坡国立大学) University of Cambridge(剑桥大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

Comments This paper has been accepted by ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12301 2025-05-20 cs.AI cs.CL 81%

Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge

Luyu Chen, Zeyu Zhang, Haoran Tan, Quanyu Dai, Hao Yang, Zhenhua Dong, Xu Chen

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments 19 pages, 3 tables, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16330 2025-05-20 cs.CL cs.AI 81%

Pruning via Merging: Compressing LLMs via Manifold Alignment Based Layer Merging

Deyuan Liu, Zhanyue Qin, Hairu Wang, Zhao Yang, Zecheng Wang, Fangying Rong, Qingbin Liu, Yanchao Hao, Xi Chen, Cunhang Fan, Zhao Lv, Zhiying Tu, Dianhui Chu, Bo Li, Dianbo Sui

机构 * Harbin Institute of Technology(哈尔滨工业大学) University of Science and Technology of China(中国科学技术大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Shandong Agricultural University(山东农业大学) Tencent Inc.(腾讯公司) Anhui University(安徽大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06267 2025-05-13 cs.SE cs.AI cs.LG 81%

AKD : Adversarial Knowledge Distillation For Large Language Models Alignment on Coding tasks

Ilyas Oulkadda, Julien Perez

机构 * Laboratoire de Recherche de l'EPITA (LRE)(EPITA研究实验室) Department of XXX, University of YYY(YYY大学XXX系) School of ZZZ, Institute of WWW(WWW研究所ZZZ学院)

专题命中 安全评测 :alignment(title);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11137 2025-05-09 cs.CL cs.AI 81%

Safety Evaluation of DeepSeek Models in Chinese Contexts

Wenjing Zhang, Xuejiao Lei, Zhaoxiang Liu, Ning Wang, Zhenhong Long, Peijun Yang, Jiaojiao Zhao, Minjie Hua, Chaoyang Ma, Kai Wang, Shiguo Lian

机构 * Unicom Data Intelligence(中国联通数据智能研究所) Data Science & Artificial Intelligence Research Institute(数据科学与人工智能研究院)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

Comments 12 pages, 2 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06172 2025-04-24 cs.AI cs.CL 81%

Multimodal Situational Safety

Kaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Anderson Compalas, Dawn Song, Xin Eric Wang

机构 * University of California, Santa Cruz(加州大学圣克ruz分校) University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

Comments ICLR 2025 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08863 2025-04-15 cs.CY cs.AI 81%

An Evaluation of Cultural Value Alignment in LLM

Nicholas Sukiennik, Chen Gao, Fengli Xu, Yong Li

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments Submitted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏