arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2508.19163 2025-08-27 cs.AI cs.HC cs.MA 57%

MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation

Ernest Lim, Yajie Vera He, Jared Joselowitz, Kate Preston, Mohita Chowdhury, Louis Williams, Aisling Higham, Katrina Mason, Mariane Melo, Tom Lawton, Yan Jia, Ibrahim Habli

机构 * Ufonia Limited(乌菲尼亚有限公司) University of York(约克大学) NHS Improvement Academy(国家健康服务改进学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 36 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19097 2025-08-27 cs.AI 57%

Reasoning LLMs in the Medical Domain: A Literature Survey

Armin Berger, Sarthak Khanna, David Berghaus, Rafet Sifa

机构 * Fraunhofer IAIS - Department of Media Engineering(弗劳恩霍夫研究所媒体工程部门) University of Bonn - Department of Computer Science(波恩大学计算机科学系) West-AI - Federal Ministry of Education and Research(西德人工智能 - 教育与研究部)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18715 2025-08-27 cs.CL 57%

EMMM, Explain Me My Model! Explainable Machine Generated Text Detection in Dialogues

Angela Yifei Yuan, Haoyi Li, Soyeon Caren Han, Christopher Leckie

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18687 2025-08-27 cs.CL 57%

Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning

Songtao Jiang, Yuxi Chen, Sibo Song, Yan Zhang, Yeying Jin, Yang Feng, Jian Wu, Zuozhu Liu

机构 * Zhejiang University, Zhejiang, China(浙江大学) Alibaba Group, Zhejiang, China(阿里巴巴集团) Angelalign Technology Inc., Shanghai, China(Angelalign技术有限公司) ChohoTech Inc., Hangzhou, China(楚合科技有限公司) Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence, Zhejiang, China(浙江省医学影像人工智能重点实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18430 2025-08-27 cs.CV cs.AI 57%

CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering

Aranya Saha, Tanvir Ahmed Khan, Ismam Nur Swapnil, Mohammad Ariful Haque

机构 * Aranya Saha ∗ , Tanvir Ahmed Khan ∗ , Ismam Nur Swapnil ∗ , and Mohammad Ariful Haque ∗(无明确机构)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 10 pages, 8 figures, Prepared for submission to IEEE Transactions on Human-Machine Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18292 2025-08-27 cs.MA cs.AI 57%

Consensus Is All You Need: Gossip-Based Reasoning Among Large Language Models

Saksham Arora

机构 * Sunnyvale, California(加州松林市)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 4 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17367 2025-08-27 cs.CV cs.AI 57%

EVM-Fusion: An Explainable Vision Mamba Architecture with Neural Algorithmic Fusion

Zichuan Yang, Yongzhi Wang

机构 * School of Mathematical Sciences, Tongji University(数学科学学院,同济大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17667 2025-08-26 cs.CV cs.AI 57%

Hierarchical Vision-Language Learning for Medical Out-of-Distribution Detection

Runhe Lai, Xinhua Lu, Kanghao Chen, Qichao Chen, Wei-Shi Zheng, Ruixuan Wang

机构 * Peng Cheng Laboratory(鹏城实验室) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Nottingham Malaysia(诺丁汉大学(马来西亚)) Key Laboratory of Machine Intelligence and Advanced Computing, MOE(机器智能与高级计算重点实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 10 pages, 2 figures, Accepted by MICCAI2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17405 2025-08-26 cs.LG cs.CR 57%

FRAME : Comprehensive Risk Assessment Framework for Adversarial Machine Learning Threats

Avishag Shapira, Simon Shigol, Asaf Shabtai

机构 * Ben-Gurion University of the Negev(贝内杰尔大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17378 2025-08-26 cs.CL 57%

UI-Level Evaluation of ALLaM 34B: Measuring an Arabic-Centric LLM via HUMAIN Chat

Omer Nacar

机构 * NAMAA Community(NAMAA社区) Riyadh - KSA(利雅得-科威特)

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09329 2025-08-26 cs.AI cs.CR 57%

When Developer Aid Becomes Security Debt: A Systematic Analysis of Insecure Behaviors in LLM Coding Agents

Matous Kozak, Roshanak Zilouchian Moghaddam, Siva Sivaraman

机构 * Microsoft(微软公司) Czech Technical University in Prague(捷克技术大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00258 2025-08-26 cs.AI 57%

Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMs

Qianqi Yan, Hongquan Li, Shan Jiang, Yang Zhao, Xinze Guan, Ching-Chen Kuo, Xin Eric Wang

机构 * University of California, Santa Cruz(加州大学圣克鲁兹分校) eBay

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06791 2025-08-26 cs.RO cs.AI cs.HC cs.MA 57%

AutoMisty: A Multi-Agent LLM Framework for Automated Code Generation in the Misty Social Robot

Xiao Wang, Lu Dong, Sahana Rangasrinivasan, Ifeoma Nwogu, Srirangaraj Setlur, Venugopal Govindaraju

机构 * State University of New York at Buffalo(纽约州立大学布法罗分校) Department of Computer Science and Engineering, Amrita School of Computing, Amrita Vishwa Vidyapeetham(计算机科学与工程系,阿米特拉学校 computing,阿米特拉世界大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted by IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01027 2025-08-26 stat.ML cs.LG 57%

Adversarial Robustness in Two-Stage Learning-to-Defer: Algorithms and Guarantees

Yannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi

机构 * School of Computing, National University of Singapore, Singapore(新加坡国立大学计算机学院) IRIT, Université de Toulouse, CNRS, Toulouse INP, Toulouse, France(图卢兹大学IRIT研究所、法国国家科学研究中心、图卢兹INP) Institute for Infocomm Research, Agency for Science, Technology and Research, Singapore(新加坡科技研究局信息与通信研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments Accepted at the 42nd International Conference on Machine Learning (ICML 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.13852 2025-08-26 cs.LG 57%

Hyperbolic Graph Neural Networks: A Review of Methods and Applications

Menglin Yang, Min Zhou, Tong Zhang, Jiahong Liu, Zhihao Li, Lujia Pan, Hui Xiong, Irwin King

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Huawei Technologies Co., Ltd.(华为技术有限公司) The Chinese University of Hong Kong(香港中文大学) Zhejiang University(浙江大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments The latest draft was circulated under the title "Hyperbolic Graph Learning: A Comprehensive Review." The present arXiv version retains the original title, "Hyperbolic Graph Neural Networks: A Review of Methods and Applications," for consistency, while incorporating substantial revisions and extensions

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16097 2025-08-25 cs.LG stat.ML 57%

Machine Learning for Medicine Must Be Interpretable, Shareable, Reproducible and Accountable by Design

Ayyüce Begüm Bektaş, Mithat Gönen

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15983 2025-08-25 cond-mat.mtrl-sci cs.LG physics.comp-ph 57%

A simulation-based training framework for machine-learning applications in ARPES

MengXing Na, Chris Zhou, Sydney K. Y. Dufresne, Matteo Michiardi, Andrea Damascelli

机构 * Quantum Matter Institute, University of British Columbia, Vancouver, British Columbia, V6T 1Z4, Canada Department of Physics \& Astronomy, University of British Columbia, Vancouver, British Columbia, V6T 1Z1, Canada

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments 9 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15926 2025-08-25 cs.CE cs.AI 57%

Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-Making

Yuanjun Feng, Vivek Choudhary, Yash Raj Shrestha

机构 * University of Lausanne(洛桑大学) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Accepted to EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17173 2025-08-25 cs.LG cs.DC 57%

Robustness of deep learning classification to adversarial input on GPUs: asynchronous parallel accumulation is a source of vulnerability

Sanjif Shanmugavelu, Mathieu Taillefumier, Christopher Culver, Vijay Ganesh, Oscar Hernandez, Ada Sedova

机构 * Maxeler Technologies, a Groq Company(Maxeler Technologies,Groq公司) ETH Zurich / CSCS(苏黎世联邦理工学院 / CSCS) Georgia Institute of Technology(佐治亚理工学院) Oak Ridge National Laboratory(橡树岭国家实验室)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14716 2025-08-22 cs.LG 57%

Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation

Tuhina Tripathi, Manya Wadhwa, Greg Durrett, Scott Niekum

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Published at COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15251 2025-08-22 eess.IV cs.AI cs.CV 57%

Explainable Knowledge Distillation for Efficient Medical Image Classification

Aqib Nazir Mir, Danish Raza Rizvi

机构 * Dept. of Computer Engineering(计算机工程系) Jamia Millia Islamia(Jamia Millia Islamia大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14927 2025-08-22 cs.GT cs.AI 57%

AI Testing Should Account for Sophisticated Strategic Behaviour

Vojtech Kovarik, Eric Olav Chen, Sami Petersen, Alexis Ghersengorin, Vincent Conitzer

机构 * Department of Computer Science(计算机科学系) Czech Technical University Prague(捷克技术大学布拉格) Global Priorities Institute(全球优先研究所) University of Oxford(牛津大学) Foundations of Cooperative AI Lab(合作人工智能基础实验室) Carnegie Mellon University(卡内基梅隆大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12259 2025-08-22 cs.CR cs.AI cs.ET 57%

Fortifying the Agentic Web: A Unified Zero-Trust Architecture Against Logic-layer Threats

Ken Huang, Yasir Mehmood, Hammad Atta, Jerry Huang, Muhammad Zeeshan Baig, Sree Bhargavi Balija

机构 * Qorvex Consulting Kleiner Perkins Wentworth Institute of Higher Education

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14654 2025-08-21 cs.AI 57%

Entropy-Constrained Strategy Optimization in Urban Floods: A Multi-Agent Framework with LLM and Knowledge Graph Integration

Peilin Ji, Xiao Xue, Simeng Wang, Wenhao Yan

机构 * College of Intelligence and Computing(智能与计算学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 17 pages including appendix, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14427 2025-08-21 cs.CL 57%

Knowledge Graph-Infused Fine-Tuning for Structured Reasoning in Large Language Models

Wuyang Zhang, Yexin Tian, Xiandong Meng, Mengjie Wang, Junliang Du

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11786 2025-08-21 stat.ML cs.LG 57%

Parallelly Tempered Generative Adversarial Nets: Toward Stabilized Gradients

Jinwon Sohn, Qifan Song

机构 * Booth School of Business, University of Chicago(芝加哥大学商学院) Department of Statistics, Purdue University(普渡大学统计学系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13465 2025-08-20 cs.AI 57%

LM Agents May Fail to Act on Their Own Risk Knowledge

Yuzhi Tang, Tianxiao Li, Elizabeth Li, Chris J. Maddison, Honghua Dong, Yangjun Ruan

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03666 2025-08-20 cs.RO cs.LG 57%

Hybrid Machine Learning Model with a Constrained Action Space for Trajectory Prediction

Alexander Fertig, Lakshman Balasubramanian, Michael Botsch

机构 * Technische Hochschule Ingolstadt, AImotion Bavaria(图腾工业大学,拜耳巴伐利亚人工智能公司) Technische Hochschule Ingolstadt, Research Center CARISSMA(图腾工业大学,CARISSMA研究中心)

专题命中 安全评测 :safety(abstract);分类 cs.LG

Comments Copyright 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works

Journal ref 2025 IEEE Intelligent Vehicles Symposium (IV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12282 2025-08-19 cs.CL cs.IR 57%

A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation

Ziyang Chen, Erxue Min, Xiang Zhao, Yunxin Li, Xin Jia, Jinzhi Liao, Jichao Li, Shuaiqiang Wang, Baotian Hu, Dawei Yin

机构 * Laboratory for Big Data and Decision, National University of Defense Technology, Changsha, China(大数据与决策实验室,国防科技大学,长沙,中国) Baidu Inc., Beijing, China(百度公司,北京,中国) Department of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), Shenzhen, China(计算机科学与技术系,哈尔滨工业大学(深圳),深圳,中国)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12100 2025-08-19 cs.AI 57%

Overcoming Knowledge Discrepancies: Structuring Reasoning Threads through Knowledge Balancing in Interactive Scenarios

Daniel Burkhardt, Xiangwei Cheng

机构 * Ferdinand Steinbeis Institute(费尔迪南·斯坦贝茨研究所)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 13 pages, 1 figure, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏