arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2502.06559 2025-05-27 cs.AI 57%

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Maria Eriksson, Erasmo Purificato, Arman Noroozian, Joao Vinagre, Guillaume Chaslot, Emilia Gomez, David Fernandez-Llorca

机构 * European Commission(欧洲委员会) Joint Research Centre (JRC)(联合研究中心)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments Under review as conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18471 2025-05-27 cs.CR cs.AI 57%

Invisible Tokens, Visible Bills: The Urgent Need to Audit Hidden Operations in Opaque LLM Services

Guoheng Sun, Ziyao Wang, Xuandong Zhao, Bowei Tian, Zheyu Shen, Yexiao He, Jinming Xing, Ang Li

机构 * University of Maryland, College Park(马里兰大学学院公园分校) University of California, Berkeley(加州大学伯克利分校) North Carolina State University(北卡罗来纳州立大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11141 2025-05-27 cs.CV cs.AI 57%

Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans

Yansheng Qiu, Li Xiao, Zhaopan Xu, Pengfei Zhou, Zheng Wang, Kaipeng Zhang

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.18815 2025-05-27 cs.LG 57%

MissionGNN: Hierarchical Multimodal GNN-based Weakly Supervised Video Anomaly Recognition with Mission-Specific Knowledge Graph Generation

Sanggeon Yun, Ryozo Masukawa, Minhyoung Na, Mohsen Imani

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Accepted to WACV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17436 2025-05-26 cs.AI 57%

Scaling Up Biomedical Vision-Language Models: Fine-Tuning, Instruction Tuning, and Multi-Modal Learning

Cheng Peng, Kai Zhang, Mengxian Lyu, Hongfang Liu, Lichao Sun, Yonghui Wu

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17154 2025-05-26 q-bio.QM cs.AI 57%

Can Large Language Models Design Biological Weapons? Evaluating Moremi Bio

Gertrude Hattoh, Jeremiah Ayensu, Nyarko Prince Ofori, Solomon Eshun, Darlington Akogo

机构 * Gertrude Hattoh(第一作者) Jeremiah Ayensu(第一作者) Nyarko Prince Ofori(第一作者) Solomon Eshun(第一作者) Darlington Akogo(第一作者)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02324 2025-05-26 cs.CY 57%

From Course to Skill: Evaluating LLM Performance in Curricular Analytics

Zhen Xu, Xinjin Li, Yingqi Huan, Veronica Minaya, Renzhe Yu

专题命中 安全评测 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16477 2025-05-23 cs.AI 57%

Advancing the Scientific Method with Large Language Models: From Hypothesis to Discovery

Yanbo Zhang, Sumeer A. Khan, Adnan Mahmud, Huck Yang, Alexander Lavin, Michael Levin, Jeremy Frey, Jared Dunnmon, James Evans, Alan Bundy, Saso Dzeroski, Jesper Tegner, Hector Zenil

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 45 pages

Journal ref npj Artificial Intelligence, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15023 2025-05-22 cs.SE cs.AI 57%

Towards a Science of Causal Interpretability in Deep Learning for Software Engineering

David N. Palacio

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments PhD thesis, To appear in ProQuest

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15372 2025-05-22 cs.CL 57%

X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System

Peng Wang, Ruihan Tao, Qiguang Chen, Mengkang Hu, Libo Qin

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments Accepted by ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14896 2025-05-22 cs.LG 57%

Feature-Weighted MMD-CORAL for Domain Adaptation in Power Transformer Fault Diagnosis

Hootan Mahmoodiyan, Maryam Ahang, Mostafa Abbasi, Homayoun Najjaran

机构 * Faculty of Engineering and Computer Science(工程与计算机科学学院) University of Victoria(维多利亚大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14402 2025-05-21 q-bio.GN cs.CL 57%

OmniGenBench: A Modular Platform for Reproducible Genomic Foundation Models Benchmarking

Heng Yang, Jack Cole, Yuan Li, Renzhi Chen, Geyong Min, Ke Li

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14352 2025-05-21 cs.LG 57%

Towards eliciting latent knowledge from LLMs with mechanistic interpretability

Bartosz Cywiński, Emil Ryd, Senthooran Rajamanoharan, Neel Nanda

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12331 2025-05-21 cs.SE cs.LG 57%

OSS-Bench: Benchmark Generator for Coding LLMs

Yuancheng Jiang, Roland Yap, Zhenkai Liang

机构 * School of Computing, National University of Singapore(计算学院,新加坡国立大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12259 2025-05-20 cs.CL 57%

Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches

Yuhang Zhou, Xutian Chen, Yixin Cao, Yuchen Ni, Yu He, Siyu Tian, Xiang Liu, Jian Zhang, Chuanjun Ji, Guangnan Ye, Xipeng Qiu

机构 * School of Computer Science, Fudan University(复旦大学计算机科学学院) Shanghai Innovation Institute(上海创新研究院) Computer Science Department, NYU Shanghai(纽约大学上海学院) DataGrand Inc.(DataGrand公司)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12201 2025-05-20 cs.CL 57%

How Reliable is Multilingual LLM-as-a-Judge?

Xiyan Fu, Wei Liu

机构 * Independent Researcher(独立研究者) Heidelberg University(海德堡大学) Heidelberg Institute for Theoretical Studies gGmbH(海德堡理论研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11829 2025-05-20 cs.CL 57%

Class Distillation with Mahalanobis Contrast: An Efficient Training Paradigm for Pragmatic Language Understanding Tasks

Chenlu Wang, Weimin Lyu, Ritwik Banerjee

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11738 2025-05-20 cs.AI 57%

Automated Real-time Assessment of Intracranial Hemorrhage Detection AI Using an Ensembled Monitoring Model (EMM)

Zhongnan Fang, Andrew Johnston, Lina Cheuy, Hye Sun Na, Magdalini Paschali, Camila Gonzalez, Bonnie A. Armstrong, Arogya Koirala, Derrick Laurel, Andrew Walker Campion, Michael Iv, Akshay S. Chaudhari, David B. Larson

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11551 2025-05-20 cs.CR cs.LG 57%

A Survey of Learning-Based Intrusion Detection Systems for In-Vehicle Network

Muzun Althunayyan, Amir Javed, Omer Rana

机构 * School of Computer Science & Informatics(计算机科学与信息学学院) Cardiff University(卡迪夫大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11535 2025-05-20 cs.RO cs.CV cs.LG 57%

Bridging Human Oversight and Black-box Driver Assistance: Vision-Language Models for Predictive Alerting in Lane Keeping Assist Systems

Yuhang Wang, Hao Zhou

机构 * Department of Civil & Environmental Engineering , University of South Florida(土木与环境工程系,佛罗里达州立大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21198 2025-05-20 cs.LG 57%

Graph Synthetic Out-of-Distribution Exposure with Large Language Models

Haoyan Xu, Zhengtao Yao, Ziyi Wang, Zhan Cheng, Xiyang Hu, Mengyuan Li, Yue Zhao

机构 * University of Southern California(南加州大学) University of Maryland, College Park(马里兰大学学院 park 分校) University of Wisconsin, Madison(威斯康星大学麦迪逊分校) Arizona State University(亚利桑那州立大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21186 2025-05-20 cs.LG 57%

GLIP-OOD: Zero-Shot Graph OOD Detection with Graph Foundation Model

Haoyan Xu, Zhengtao Yao, Xuzhi Zhang, Ziyi Wang, Langzhou He, Yushun Dong, Philip S. Yu, Mengyuan Li, Yue Zhao

机构 * University of Southern California(南加州大学) University of Maryland, College Park(马里兰大学学院公园分校) University of Illinois Chicago(伊利诺伊大学芝加哥分校) Florida State University(佛罗里达州立大学)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02564 2025-05-20 cs.CV cs.AI q-bio.NC 57%

Probing Human Visual Robustness with Neurally-Guided Deep Neural Networks

Zhenan Shao, Linjian Ma, Yiqing Zhou, Yibo Jacky Zhang, Sanmi Koyejo, Bo Li, Diane M. Beck

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Cornell University(康奈尔大学) Stanford University(斯坦福大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11342 2025-05-19 cs.LG math.OC 57%

Sobolev Training of End-to-End Optimization Proxies

Andrew W. Rosemberg, Joaquim Dias Garcia, Russell Bent, Pascal Van Hentenryck

机构 * Georgia Institute of Technology(佐治亚理工学院) Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室) PSR

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments 9 Pages, 4 Figures, 5 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10609 2025-05-19 cs.CR cs.AI cs.MA cs.NI 57%

Agent Name Service (ANS): A Universal Directory for Secure AI Agent Discovery and Interoperability

Ken Huang, Vineeth Sai Narajala, Idan Habler, Akram Sheriff

机构 * Amazon Web Services(亚马逊网络服务) Intuit Cisco Systems(Cisco系统)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 15 pages, 6 figures, 6 code listings, Supported and endorsed by OWASP GenAI ASI Project

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10409 2025-05-16 cs.CL 57%

Are LLM-generated plain language summaries truly understandable? A large-scale crowdsourced evaluation

Yue Guo, Jae Ho Sohn, Gondy Leroy, Trevor Cohen

机构 * School of Information Sciences, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校信息科学学院) Department of Radiology and Biomedical Imaging, University of California, San Francisco(加州大学旧金山分校放射科与生物医学成像系) Eller College of Management, University of Arizona(亚利桑那大学埃勒管理学院) Biomedical Informatics and Medical Education, University of Washington(华盛顿大学生物医学信息学与医学教育)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09806 2025-05-16 cs.HC cs.CY 57%

Learn, Explore and Reflect by Chatting: Understanding the Value of an LLM-Based Voting Advice Application Chatbot

Jianlong Zhu, Manon Kempermann, Vikram Kamath Cannanure, Alexander Hartland, Rosa M. Navarrete, Giuseppe Carteny, Daniela Braun, Ingmar Weber

专题命中 安全评测 :trustworthy(abstract);分类 cs.CY

Comments Accepted to ACM CUI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08604 2025-05-16 cs.RO cs.AI 57%

EMMOE: A Comprehensive Benchmark for Embodied Mobile Manipulation in Open Environments

Dongping Li, Tielong Cai, Tianci Tang, Wenhao Chai, Katherine Rose Driggs-Campbell, Gaoang Wang

机构 * Zhejiang University(浙江大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Washington(华盛顿大学)

专题命中 安全评测 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08775 2025-05-14 cs.CL 57%

HealthBench: Evaluating Large Language Models Towards Improved Human Health

Rahul K. Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero-Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, Johannes Heidecke, Karan Singhal

机构 * OpenAI

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments Blog: https://openai.com/index/healthbench/ Code: https://github.com/openai/simple-evals

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16254 2025-05-14 cs.CR cs.CL 57%

Adversarial Robustness through Dynamic Ensemble Learning

Hetvi Waghela, Jaydip Sen, Sneha Rakshit

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments This is the accepted version of our paper for the 2024 IEEE Silchar Subsection Conference (IEEE SILCON24), held from November 15 to 17, 2024, at the National Institute of Technology (NIT), Agartala, India. The paper is 6 pages long and contains 3 Figures and 7 Tables

Journal ref 2024 IEEE Silchar Subsection Conference (SILCON 2024) Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏