arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-01 至 2025-10-01 共收录 19 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 19 篇

2501.15463 2025-10-01 cs.HC cs.AI cs.CL 81%

Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?

Hua Shen, Nicholas Clark, Tanushree Mitra

机构 * University of Washington(华盛顿大学) NYU Shanghai(纽约大学上海校区) New York University(纽约大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 Main Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26167 2025-10-01 cs.AI 79%

'Too much alignment; not enough culture': Re-balancing cultural alignment practices in LLMs

Eric J. W. Orlowski, Hakim Norhashim, Tristan Koh Ly Wey

机构 * AI Singapore, National University of Singapore(AI新加坡,国立新加坡大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments 8 pages, no figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25885 2025-10-01 cs.AI 79%

SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents

Ruolin Chen, Yinqian Sun, Jihang Wang, Mingyang Lv, Qian Zhang, Yi Zeng

机构 * Brain-inspired Cognitive AI Lab, Institute of Automation, Chinese Academy of Sciences(脑启发认知AI实验室,自动化研究所,中国科学院) Beijing Key Laboratory of Safe AI and Superalignment(北京安全AI与超对齐关键实验室) Beijing Institute of AI Safety and Governance(北京AI安全与治理研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Long-term AI(长期AI)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26122 2025-10-01 math.NA cs.NA 78%

Trustworthy AI in numerics: On verification algorithms for neural network-based PDE solvers

Emil Haugen, Alexei Stepanenko, Anders C. Hansen

专题命中 安全评测 :trustworthy(title,abstract)

Comments 25 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04974 2025-10-01 eess.AS cs.SD 78%

From Voice to Safety: Language AI Powered Pilot-ATC Communication Understanding for Airport Surface Movement Collision Risk Assessment

Yutian Pang, Andrew Paul Kendall, Alex Porcayo, Mariah Barsotti, Anahita Jain, John-Paul Clarke

机构 * Department of Aerospace Engineering and Engineering Mechanics(航空航天工程与工程力学系)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26417 2025-10-01 cs.AI cs.LG 62%

OntoAligner Meets Knowledge Graph Embedding Aligners

Hamed Babaei Giglou, Jennifer D'Souza, Sören Auer, Mahsa Sanaei

机构 * TIB -- Leibniz Information Centre for Science and Technology(莱比锡信息科学与技术研究中心) University of Tabriz(塔布里兹大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 10 pages of main content, 3 page references, 3 figures. Accepted to Ontology Matching Workshop at ISWC

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22048 2025-10-01 cs.CL cs.LG 62%

ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models

Chung-En Sun, Ge Yan, Tsui-Wei Weng

机构 * UCSD CSE(加州大学圣地亚哥分校计算机科学系) UCSD HDSI(加州大学圣地亚哥分校人类发展与应用科学学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11843 2025-10-01 cs.CL cs.AI 62%

Preemptive Detection and Correction of Misaligned Actions in LLM Agents

Haishuo Fang, Xiaodan Zhu, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab, Department of Computer Science, TU Darmstadt(图宾根大学计算机科学系) National Research Center for Applied Cybersecurity ATHENE, Germany(德国应用网络安全国家研究中心ATHENE) Department of Electrical and Computer Engineering & Ingenuity Labs Research Institute, Queen’s University, Canada(加拿大女王大学电气与计算机工程系及创新实验室研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26640 2025-10-01 cs.LG cs.CR 57%

SPATA: Systematic Pattern Analysis for Detailed and Transparent Data Cards

João Vitorino, Eva Maia, Isabel Praça, Carlos Soares

机构 * GECAD, ISEP, Polytechnic of Porto Faculty of Engineering, University of Porto(GECAD、ISEP、波尔图理工学院工程学院、波尔图大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 16 pages, 3 tables, 6 figures, SynDAiTE, ECML PKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26636 2025-10-01 cs.LG 57%

AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond

Shangding Gu, Xiaohan Wang, Donghao Ying, Haoyu Zhao, Runing Yang, Ming Jin, Boyi Li, Marco Pavone, Serena Yeung-Levy, Jun Wang, Dawn Song, Costas Spanos

机构 * UC Berkeley(伯克利大学) Stanford(斯坦福大学) UCL(伦敦大学学院) Virginia Tech(弗吉尼亚理工学院) Nvidia(英伟达公司)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26276 2025-10-01 cs.CL cs.SD 57%

Optimizing Speech Language Models for Acoustic Consistency

Morteza Rohanian, Michael Krauthammer

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26225 2025-10-01 cs.CV cs.AI 57%

An Experimental Study on Generating Plausible Textual Explanations for Video Summarization

Thomas Eleftheriadis, Evlampios Apostolidis, Vasileios Mezaris

机构 * IEEE CBMI 2025

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments IEEE CBMI 2025. This is the authors' accepted version. The final publication is available at https://ieeexplore.ieee.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25651 2025-10-01 cs.AI 57%

AutoLabs: Cognitive Multi-Agent Systems with Self-Correction for Autonomous Chemical Experimentation

Gihan Panapitiya, Emily Saldanha, Heather Job, Olivia Hess

机构 * Pacific Northwest National Laboratory(太平洋西北国家实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25637 2025-10-01 cs.LG 57%

How Does Preconditioning Guide Feature Learning in Deep Neural Networks?

Kotaro Yoshida, Atsushi Nitanda

机构 * Institute of Science Tokyo(东京科学研究院) Agency for Science, Technology and Research (A*STAR)(科技研究局) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20411 2025-10-01 cs.CR cs.AI 57%

Adversarial Defense in Cybersecurity: A Systematic Review of GANs for Threat Detection and Mitigation

Tharcisse Ndayipfukamiye, Jianguo Ding, Doreen Sebastian Sarwatt, Adamu Gaston Philipo, Huansheng Ning

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 36 pages, 10 tables, 4figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23326 2025-10-01 cs.HC cs.CY 57%

Designing the Future of Entrepreneurship Education: Exploring an AI-Empowered Scaffold System for Business Plan Development

Junhua Zhu, Lan Luo

专题命中 安全评测 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.09349 2025-10-01 cs.AI cs.HC cs.MA 57%

Establishing Shared Query Understanding in an Open Multi-Agent System

Nikolaos Kondylidis, Ilaria Tiddi, Annette ten Teije

机构 * Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 9 pages. International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2023), London, United Kingdom

Journal ref In Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems (AAMAS 23). International Foundation for Autonomous Agents and Multiagent Systems, Richland

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26223 2025-10-01 q-bio.GN 50%

Nephrobase Cell+: Multimodal Single-Cell Foundation Model for Decoding Kidney Biology

Chenyu Li, Elias Ziyadeh, Yash Sharma, Bernhard Dumoulin, Jonathan Levinsohn, Eunji Ha, Siyu Pan, Vishwanatha Rao, Madhav Subramaniyam, Mario Szegedy, Nancy Zhang, Katalin Susztak

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25974 2025-10-01 cs.NI cs.MA 50%

OpenID Connect for Agents (OIDC-A) 1.0: A Standard Extension for LLM-Based Agent Identity and Authorization

Subramanya Nagabhushanaradhya

专题命中 安全评测 :trustworthy(abstract)

Comments 10 pages, 5 tables, 2 code listings. Specification proposal available at https://github.com/subramanya1997/oidc-a/

详情

展开后加载摘要…

URL PDF HTML 收藏