arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-10-01 至 2025-10-01 共收录 57 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 19 篇

2509.26417 2025-10-01 cs.AI cs.LG 62%

OntoAligner Meets Knowledge Graph Embedding Aligners

Hamed Babaei Giglou, Jennifer D'Souza, Sören Auer, Mahsa Sanaei

机构 * TIB -- Leibniz Information Centre for Science and Technology(莱比锡信息科学与技术研究中心) University of Tabriz(塔布里兹大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 10 pages of main content, 3 page references, 3 figures. Accepted to Ontology Matching Workshop at ISWC

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22048 2025-10-01 cs.CL cs.LG 62%

ThinkEdit: Interpretable Weight Editing to Mitigate Overly Short Thinking in Reasoning Models

Chung-En Sun, Ge Yan, Tsui-Wei Weng

机构 * UCSD CSE(加州大学圣地亚哥分校计算机科学系) UCSD HDSI(加州大学圣地亚哥分校人类发展与应用科学学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.LG

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11843 2025-10-01 cs.CL cs.AI 62%

Preemptive Detection and Correction of Misaligned Actions in LLM Agents

Haishuo Fang, Xiaodan Zhu, Iryna Gurevych

机构 * Ubiquitous Knowledge Processing Lab, Department of Computer Science, TU Darmstadt(图宾根大学计算机科学系) National Research Center for Applied Cybersecurity ATHENE, Germany(德国应用网络安全国家研究中心ATHENE) Department of Electrical and Computer Engineering & Ingenuity Labs Research Institute, Queen’s University, Canada(加拿大女王大学电气与计算机工程系及创新实验室研究院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted by EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26640 2025-10-01 cs.LG cs.CR 57%

SPATA: Systematic Pattern Analysis for Detailed and Transparent Data Cards

João Vitorino, Eva Maia, Isabel Praça, Carlos Soares

机构 * GECAD, ISEP, Polytechnic of Porto Faculty of Engineering, University of Porto(GECAD、ISEP、波尔图理工学院工程学院、波尔图大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 16 pages, 3 tables, 6 figures, SynDAiTE, ECML PKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26636 2025-10-01 cs.LG 57%

AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond

Shangding Gu, Xiaohan Wang, Donghao Ying, Haoyu Zhao, Runing Yang, Ming Jin, Boyi Li, Marco Pavone, Serena Yeung-Levy, Jun Wang, Dawn Song, Costas Spanos

机构 * UC Berkeley(伯克利大学) Stanford(斯坦福大学) UCL(伦敦大学学院) Virginia Tech(弗吉尼亚理工学院) Nvidia(英伟达公司)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26276 2025-10-01 cs.CL cs.SD 57%

Optimizing Speech Language Models for Acoustic Consistency

Morteza Rohanian, Michael Krauthammer

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26225 2025-10-01 cs.CV cs.AI 57%

An Experimental Study on Generating Plausible Textual Explanations for Video Summarization

Thomas Eleftheriadis, Evlampios Apostolidis, Vasileios Mezaris

机构 * IEEE CBMI 2025

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments IEEE CBMI 2025. This is the authors' accepted version. The final publication is available at https://ieeexplore.ieee.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25651 2025-10-01 cs.AI 57%

AutoLabs: Cognitive Multi-Agent Systems with Self-Correction for Autonomous Chemical Experimentation

Gihan Panapitiya, Emily Saldanha, Heather Job, Olivia Hess

机构 * Pacific Northwest National Laboratory(太平洋西北国家实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25637 2025-10-01 cs.LG 57%

How Does Preconditioning Guide Feature Learning in Deep Neural Networks?

Kotaro Yoshida, Atsushi Nitanda

机构 * Institute of Science Tokyo(东京科学研究院) Agency for Science, Technology and Research (A*STAR)(科技研究局) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20411 2025-10-01 cs.CR cs.AI 57%

Adversarial Defense in Cybersecurity: A Systematic Review of GANs for Threat Detection and Mitigation

Tharcisse Ndayipfukamiye, Jianguo Ding, Doreen Sebastian Sarwatt, Adamu Gaston Philipo, Huansheng Ning

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 36 pages, 10 tables, 4figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23326 2025-10-01 cs.HC cs.CY 57%

Designing the Future of Entrepreneurship Education: Exploring an AI-Empowered Scaffold System for Business Plan Development

Junhua Zhu, Lan Luo

专题命中 安全评测 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.09349 2025-10-01 cs.AI cs.HC cs.MA 57%

Establishing Shared Query Understanding in an Open Multi-Agent System

Nikolaos Kondylidis, Ilaria Tiddi, Annette ten Teije

机构 * Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 9 pages. International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2023), London, United Kingdom

Journal ref In Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems (AAMAS 23). International Foundation for Autonomous Agents and Multiagent Systems, Richland

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26223 2025-10-01 q-bio.GN 50%

Nephrobase Cell+: Multimodal Single-Cell Foundation Model for Decoding Kidney Biology

Chenyu Li, Elias Ziyadeh, Yash Sharma, Bernhard Dumoulin, Jonathan Levinsohn, Eunji Ha, Siyu Pan, Vishwanatha Rao, Madhav Subramaniyam, Mario Szegedy, Nancy Zhang, Katalin Susztak

专题命中 安全评测 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25974 2025-10-01 cs.NI cs.MA 50%

OpenID Connect for Agents (OIDC-A) 1.0: A Standard Extension for LLM-Based Agent Identity and Authorization

Subramanya Nagabhushanaradhya

专题命中 安全评测 :trustworthy(abstract)

Comments 10 pages, 5 tables, 2 code listings. Specification proposal available at https://github.com/subramanya1997/oidc-a/

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 4 篇

2507.14688 2025-10-01 cs.CL cs.AI cs.LG 75%

Mind the Gap: A Review of Arabic Post-Training Datasets and Their Limitations

Mohammed Alkhowaiter, Norah Alshahrani, Saied Alshahrani, Reem I. Masoud, Alaa Alzahrani, Deema Alnuhait, Emad A. Alghamdi, Khalid Almubarak

机构 * Refine AI ASAS AI University of Bisha(比沙大学) University College London(伦敦大学学院) King Salman Global Academy for Arabic(萨勒曼全球阿拉伯学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) King Abdulaziz University(阿卜杜勒阿齐兹大学) HUMAIN

专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26584 2025-10-01 cs.AI cs.IR cs.LG cs.SE 62%

Fairness Testing in Retrieval-Augmented Generation: How Small Perturbations Reveal Bias in Small Language Models

Matheus Vinicius da Silva de Oliveira, Jonathan de Andrade Silva, Awdren de Lima Fontao

机构 * Faculty of Computing - Federal University of Mato Grosso do Sul(计算机学院 - 短暂戈亚那联邦大学)

专题命中 AI治理与伦理 :prompt injection(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07931 2025-10-01 cs.CY cs.AI 62%

Educating a Responsible AI Workforce: Piloting a Curricular Module on AI Policy in a Graduate Machine Learning Course

James Weichert, Hoda Eldardiry

机构 * Department of Computer Science Virginia Tech(计算机科学系弗吉尼亚理工大学)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

Comments Accepted at 2025 ASEE Annual Conference & Exposition

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25963 2025-10-01 cs.CV 50%

Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation

Longzhen Yang, Zhangkai Ni, Ying Wen, Yihang Liu, Lianghua He, Heng Tao Shen

机构 * Tongji University(同济大学) East China Normal University(华东师范大学) Shanghai Eye Disease Prevention and Treatment Center(上海眼病防治中心)

专题命中 AI治理与伦理 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 9 篇

2509.25253 2025-10-01 cs.LG cs.AI 81%

Knowledge distillation through geometry-aware representational alignment

Prajjwal Bhattarai, Mohammad Amjad, Dmytro Zhylko, Tuka Alhanai

机构 * New York University Abu Dhabi(纽约大学阿布扎赫德分校) New York University(纽约大学) Tandon School of Engineering(Tandon工程学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19607 2025-10-01 cs.HC cs.AI 79%

Enabling Rapid Shared Human-AI Mental Model Alignment via the After-Action Review

Edward Gu, Ho Chit Siu, Melanie Platt, Isabelle Hurley, Jaime Peña, Rohan Paleja

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Accepted to the Cooperative Multi-Agent Systems Decision-making and Learning:Human-Multi-Agent Cognitive Fusion Workshop at AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26625 2025-10-01 cs.LG cs.AI cs.CV cs.MM 62%

Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training

Junlin Han, Shengbang Tong, David Fan, Yufan Ren, Koustuv Sinha, Philip Torr, Filippos Kokkinos

机构 * Meta Superintelligence Labs(Meta 超智能实验室) University of Oxford(牛津大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Project page: https://junlinhan.github.io/projects/lsbs/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26239 2025-10-01 cs.LG cs.AI stat.ML 62%

Sandbagging in a Simple Survival Bandit Problem

Joel Dyer, Daniel Jarne Ornia, Nicholas Bishop, Anisoara Calinescu, Michael Wooldridge

机构 * University of Oxford(牛津大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments Forthcoming in the "Reliable ML from Unreliable Data Workshop" at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25220 2025-10-01 cs.CL cs.LG 62%

Cyclic Ablation: Testing Concept Localization against Functional Regeneration in AI

Eduard Kapelko

机构 * Eduard Kapelko(独立研究者)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.LG

Comments Code is available at: https://www.kaggle.com/code/kapedalex/cycleablationpublic/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26128 2025-10-01 cs.AI 57%

MEDAKA: Construction of Biomedical Knowledge Graphs Using Large Language Models

Asmita Sengupta, David Antony Selby, Sebastian Josef Vollmer, Gerrit Großmann

机构 * Department of Data Science and its Applications, German Research Center for Artificial Intelligence (DFKI GmbH)(数据科学及其应用系,德国人工智能研究中心(DFKI GmbH)) Department of Computer Science, University of Kaiserslautern–Landau (RPTU)(计算机科学系,凯撒斯劳滕-兰道大学(RPTU))

专题命中 其他安全 :safety(abstract);分类 cs.AI

Comments 9 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25243 2025-10-01 cs.SE cs.AI 57%

Reinforcement Learning-Guided Chain-of-Draft for Token-Efficient Code Generation

Xunzhu Tang, Iyiola Emmanuel Olatunji, Tiezhu Sun, Jacques Klein, Tegawende F. Bissyande

机构 * University of Luxembourg(卢森堡大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26012 2025-10-01 cs.CV 50%

SETR: A Two-Stage Semantic-Enhanced Framework for Zero-Shot Composed Image Retrieval

Yuqi Xiao, Yingying Zhu

机构 * Yuqi Xiao, Yingying Zhu

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25728 2025-10-01 cond-mat.mtrl-sci physics.app-ph 50%

Fingerprinting Organic Molecules for the Inverse Design of Two-Dimensional Hybrid Perovskites with Target Energetics

Yongxin Lyu, Yifan Zhou, Yu Zhang, Yang Yang, Bosen Zou, Qiang Weng, Tong Xie, Claudio Cazorla, Jianhua Hao, Jun Yin, Tom Wu

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏