arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9346 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9346 篇

2510.07575 2025-10-16 cs.AI cs.LG 62%

Benchmarking is Broken -- Don't Let AI be its Own Judge

Zerui Cheng, Stella Wohnig, Ruchika Gupta, Samiul Alam, Tassallah Abdullahi, João Alves Ribeiro, Christian Nielsen-Garcia, Saif Mir, Siran Li, Jason Orender, Seyed Ali Bahrainian, Daniel Kirste, Aaron Gokaslan, Mikołaj Glinka, Carsten Eickhoff, Ruben Wolff

机构 * Princeton University(普林斯顿大学) CISPA Helmholtz Center for Information Security(CISPA海德堡信息安全中心) Michigan State University(密歇根州立大学) Ohio State University(俄亥俄州立大学) Brown University(布朗大学) Massachusetts Institute of Technology(麻省理工学院) University of California, Los Angeles(加州大学洛杉矶分校) University of Tübingen(图宾根大学) Old Dominion University(旧 Dominion 大学) Technical University of Munich(慕尼黑技术大学) Cornell University(康奈尔大学) Forest AI(森林AI)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments 14 pages; Accepted to NeurIPS 2025. Link to poster: https://neurips.cc/virtual/2025/poster/121919; Link to project website: https://www.peerbench.ai/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02594 2025-10-16 cs.HC cs.AI cs.CL 62%

A Risk Taxonomy and Reflection Tool for Large Language Model Adoption in Public Health

Jiawei Zhou, Amy Z. Chen, Darshi Shah, Laura M. Schwab Reese, Munmun De Choudhury

机构 * Georgia Institute of Technology(佐治亚理工学院) Purdue University(普渡大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23042 2025-10-14 cs.CV cs.AI cs.LG cs.MM cs.RO 62%

Goal-Based Vision-Language Driving

Santosh Patapati, Trisanth Srinivasan

机构 * Dept. of HCI(人机交互系) Cyrion Labs(Cyrion实验室)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19981 2025-10-14 cs.LG cs.CL 62%

Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets

Adam Younsi, Ahmed Attia, Abdalgader Abubaker, Mohamed El Amine Seddik, Hakim Hacid, Salem Lahlou

机构 * Technology Innovation Institute(技术创新研究所) Mohamed Bin Zayed University of Artificial Intelligence(莫罕默德·本·扎耶德人工智能大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09479 2025-10-14 cs.AI cs.CL 62%

Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation

Zhiqing Cui, Jiahao Yuan, Hanqing Wang, Yanshu Li, Chenxu Du, Zhenglong Ding

机构 * Nanjing University of Information Science \& Technology Nanjing China East China Normal University Shanghai China The Hong Kong University of Science Brown University Providence America Southwest Jiaotong University Chengdu China Nanjing University of Information Science \& Technology East China Normal University Brown University Southwest Jiaotong University

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 5 figures, accepted to appear in the Proceedings of the 33rd ACM International Conference on Multimedia (MM '25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09738 2025-10-14 cs.CL cs.AI 62%

Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement

Steve Han, Gilberto Titericz Junior, Tom Balough, Wenfei Zhou

机构 * NVIDIA Corporation(NVIDIA公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 10 pages, 1 figure, 4 tables, under review as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10983 2025-10-14 cs.LG cs.AI cs.CR 62%

GenoArmory: A Unified Evaluation Framework for Adversarial Attacks on Genomic Foundation Models

Haozheng Luo, Chenghao Qiu, Yimin Wang, Shang Wu, Jiahao Yu, Zhenyu Pan, Weian Mao, Haoyang Fang, Hao Xu, Han Liu, Binghui Wang, Yan Chen

机构 * Department of Computer Science, Northwestern University(西北大学计算机科学系) Department of Computer Science and Engineering, Texas A&M University(德克萨斯农工大学计算机科学与工程系) Department of Computer Science, Illinois Institute of Technology(伊利诺伊理工学院计算机科学系) Department of Statistics and Data Science, Northwestern University(西北大学统计与数据科学系) Department of Computer Science and Engineering, University of Michigan(密歇根大学计算机科学与工程系) Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology(麻省理工学院电气工程与计算机科学系) Department of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学系) Department of Medicine at Brigham and Women’s Hospital and Harvard Medical School(布里格姆和妇女医院及哈佛医学院医学系)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14858 2025-10-14 cs.AI cs.CL 62%

Retrieval is Not Enough: Enhancing RAG Reasoning through Test-Time Critique and Optimization

Jiaqi Wei, Hao Zhou, Xiang Zhang, Di Zhang, Zijie Qiu, Wei Wei, Jinzhe Li, Wanli Ouyang, Siqi Sun

机构 * Zhejiang University(浙江大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) South China University of Technology(华南理工大学) University of British Columbia(不列颠哥伦比亚大学) Fudan University(复旦大学) University of Hong Kong(香港大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08114 2025-10-10 cs.AI cs.CL 62%

Can Risk-taking AI-Assistants suitably represent entities

Ali Mazyaki, Mohammad Naghizadeh, Samaneh Ranjkhah Zonouzaghi, Amirhossein Farshi Sotoudeh

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07626 2025-10-10 cs.LG cs.CL 62%

LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics

Chongyu Fan, Changsheng Wang, Yancheng Huang, Soumyadeep Pal, Sijia Liu

机构 * Michigan State University(密歇根州立大学) IBM Research(IBM研究院)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04363 2025-10-10 cs.SE cs.AI cs.CL 62%

MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models

Hyunjun Kim, Sejong Kim

机构 * KAIST(韩国科学技术院)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2025 Workshop on Lock-LLM

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25256 2025-10-10 cs.CY cs.AI 62%

The Sandbox Configurator: A Framework to Support Technical Assessment in AI Regulatory Sandboxes

Alessio Buscemi, Thibault Simonetto, Daniele Pagani, German Castignani, Maxime Cordy, Jordi Cabot

机构 * Luxembourg Institute of Science and Technology (LIST)(卢森堡科学与技术研究院) University of Luxembourg(卢森堡大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17694 2025-10-10 cs.CL cs.AI 62%

Evaluating LLM-Generated Versus Human-Authored Responses in Role-Play Dialogues

Dongxu Lu, Johan Jeuring, Albert Gatt

机构 * Utrecht University(乌得勒支大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted for publication at the 18th International Natural Language Generation Conference (INLG 2025). Revised version: improved image quality and minor corrections. No change to conclusions

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00529 2025-10-10 cs.CL cs.CY 62%

Modeling Motivated Reasoning in Law: Evaluating Strategic Role Conditioning in LLM Summarization

Eunjung Cho, Alexander Hoyle, Yoan Hermstrüwer

机构 * ETH Zurich(苏黎世联邦理工学院) University of Zurich(苏黎世大学) Max Planck Institute for Research on Collective Goods(集体利益研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.CY

Comments Accepted at NLLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17717 2025-10-10 cs.CL cs.AI 62%

From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes

Karen Zhou, John Giorgi, Pranav Mani, Peng Xu, Davis Liang, Chenhao Tan

机构 * University of Chicago(芝加哥大学) Abridge

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08713 2025-10-10 cs.LG cs.AI 62%

ProtoECGNet: Case-Based Interpretable Deep Learning for Multi-Label ECG Classification with Contrastive Learning

Sahil Sethi, David Chen, Thomas Statchen, Michael C. Burkhart, Nipun Bhandari, Bashar Ramadan, Brett Beaulieu-Jones

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments Accepted to PMLR 298, 10th Machine Learning for Healthcare Conference (MLHC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.02944 2025-10-10 cs.LG cs.AI cs.SY eess.SY 62%

Foundation Models for Structural Health Monitoring

Luca Benfenati, Daniele Jahier Pagliari, Luca Zanatta, Yhorman Alexander Bedoya Velez, Andrea Acquaviva, Massimo Poncino, Enrico Macii, Luca Benini, Alessio Burrello

机构 * DAUIN, Politecnico di Torino(达乌因,托斯卡纳理工学院) DEI, University of Bologna(电子工程学院,博洛尼亚大学) DIST, Politecnico di Torino(信息与通信技术学院,托斯卡纳理工学院) D-ITET, ETH Zurich(信息与通信技术系,苏黎世联邦理工学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 17 pages, 6 tables, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07243 2025-10-09 cs.CL cs.AI 62%

LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation

Joseph Enguehard, Morgane Van Ermengem, Kate Atkinson, Sujeong Cha, Arijit Ghosh Chowdhury, Prashanth Kallur Ramaswamy, Jeremy Roghair, Hannah R Marlowe, Carina Suzana Negreanu, Kitty Boxall, Diana Mincu

机构 * Robin AI Amazon Web Services(亚马逊网络服务)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

Comments Published in Natural Legal Language Processing - EMNLP Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03487 2025-10-09 cs.LG cs.AI cs.CR q-bio.BM q-bio.QM 62%

SafeProtein: Red-Teaming Framework and Benchmark for Protein Foundation Models

Jigang Fan, Zhenghong Zhou, Ruofan Jin, Le Cong, Mengdi Wang, Zaixi Zhang

机构 * Peking University(北京大学) Stanford University(斯坦福大学) Princeton University(普林斯顿大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06253 2025-10-09 cs.CY cs.AI 62%

LLM-Driven Rubric-Based Assessment of Algebraic Competence in Multi-Stage Block Coding Tasks with Design and Field Evaluation

Yong Oh Lee, Byeonghun Bang, Sejun Oh

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05976 2025-10-08 cs.CV cs.AI cs.LG 62%

Diffusion Models for Low-Light Image Enhancement: A Multi-Perspective Taxonomy and Performance Analysis

Eashan Adhikarla, Yixin Liu, Brian D. Davison

机构 * Lehigh University(莱维大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05972 2025-10-08 cs.CL cs.AI 62%

LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language

Periklis Mantenoglou, Rishi Hazra, Pedro Zuidberg Dos Martires, Luc De Raedt

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05310 2025-10-08 cs.CL cs.AI 62%

RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts

Yining She, Daniel W. Peterson, Marianne Menglin Liu, Vikas Upadhyay, Mohammad Hossein Chaghazardi, Eunsuk Kang, Dan Roth

机构 * Carnegie Mellon University(卡内基梅隆大学) Oracle Cloud Infrastructure(Oracle 云基础设施) University of Pennsylvania(宾夕法尼亚大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12543 2025-10-08 cs.AI cs.CV cs.LG 62%

Human + AI for Accelerating Ad Localization Evaluation

Harshit Rajgarhia, Shivali Dalmia, Mengyang Zhao, Mukherji Abhishek, Kiran Ganesh

机构 * Centific Global Solutions Inc.(Centific全球解决方案公司)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11676 2025-10-08 cs.LG cs.AI stat.ME stat.ML 62%

SKADA-Bench: Benchmarking Unsupervised Domain Adaptation Methods with Realistic Validation On Diverse Modalities

Yanis Lalou, Théo Gnassounou, Antoine Collas, Antoine de Mathelin, Oleksii Kachaiev, Ambroise Odonnat, Alexandre Gramfort, Thomas Moreau, Rémi Flamary

机构 * École Polytechnique, IP Paris, CMAP, UMR 7641(巴黎理工学院) Université Paris-Saclay, Inria, CEA(巴黎萨克雷大学) Inria(法国国家信息与自动化技术研究院) CEA(法国原子能机构) Centre Borelli, ENS Paris-Saclay(巴黎-萨克雷大学博雷利中心) Università degli Studi di Genova(热那亚大学) Inria, Univ. Rennes 2, CNRS, IRISA(法国国家信息与自动化技术研究院、里昂二大学、法国国家科学研究中心、IRISA)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Published in Transactions on Machine Learning Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01238 2025-10-07 cs.CL cs.LG 62%

Silent Tokens, Loud Effects: Padding in LLMs

Rom Himelstein, Amit LeVi, Yonatan Belinkov, Avi Mendelson

机构 * Department of Data and Decision Science, Technion - Israel Institute of Technology(数据与决策科学系,技术离子理工学院) Department of Computer Science, Technion - Israel Institute of Technology(计算机科学系,技术离子理工学院)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.LG

Comments Accepted to NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00476 2025-10-06 cs.SE cs.AI cs.LG 62%

Analyzing Latent Concepts in Code Language Models

Arushi Sharma, Vedant Pungliya, Christopher J. Quinn, Ali Jannesari

机构 * Iowa State University(爱荷华州立大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01792 2025-10-03 cs.CL cs.AI cs.IR 62%

Comparison of Unsupervised Metrics for Evaluating Judicial Decision Extraction

Ivan Leonidovich Litvak, Anton Kostin, Fedor Lashkin, Tatiana Maksiyan, Sergey Lagutin

机构 * Moscow Center for Advanced Studies(莫斯科高级研究学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments 28 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01270 2025-10-03 cs.CL cs.AI 62%

Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection

Hoang Phan, Victor Li, Qi Lei

机构 * New York University(纽约大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25264 2025-10-03 cs.DB cs.AI cs.LG cs.SE 62%

GeoSQL-Eval: First Evaluation of LLMs on PostGIS-Based NL2GeoSQL Queries

Shuyang Hou, Haoyue Jiao, Ziqi Liu, Lutong Xie, Guanyu Chen, Shaowen Wu, Xuefeng Guan, Huayi Wu

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏