arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9434 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9434 篇

2510.01295 2025-10-03 cs.AI cs.MA 57%

The Social Laboratory: A Psychometric Framework for Multi-Agent LLM Evaluation

Zarreen Reza

机构 * Independent researcher(独立研究者)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26301 2025-10-03 cs.LG cs.HC 57%

NeuroTTT: Bridging Pretraining-Downstream Task Misalignment in EEG Foundation Models via Test-Time Training

Suli Wang, Yangshen Deng, Zhenghua Bao, Xinyu Zhan, Yiqun Duan

机构 * Technical University of Darmstadt(达姆斯塔特技术大学) University of Edinburgh(爱丁堡大学) University of Technology Sydney(悉尼技术大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21285 2025-10-03 cs.CL 57%

Double-Checker: Enhancing Reasoning of Slow-Thinking LLMs via Self-Critical Fine-Tuning

Xin Xu, Tianhao Chen, Fan Zhang, Wanlong Liu, Pengxiang Li, Ajay Kumar Jaiswal, Yuchen Yan, Jishan Hu, Yang Wang, Hao Chen, Shiwei Liu, Shizhe Diao, Can Yang, Lu Yin

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) University of Electronic Science and Technology of China(电子科技大学) Dalian University of Technology(大连理工大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) Zhejiang University(浙江大学) University of Oxford(牛津大学) NVIDIA(NVIDIA公司) University of Surrey(萨里大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02064 2025-10-03 cs.CY cs.HC 57%

The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims

Kiana Jafari Meimandi, Gabriela Aránguiz-Dias, Grace Ra Kim, Lana Saadeddin, Allie Griffith, Mykel J. Kochenderfer

专题命中 安全评测 :safety(abstract);分类 cs.CY

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08198 2025-10-03 stat.ML cs.LG 57%

SIM-Shapley: A Stable and Computationally Efficient Approach to Shapley Value Approximation

Wangxuan Fan, Siqi Li, Doudou Zhou, Yohei Okada, Chuan Hong, Molei Liu, Nan Liu

机构 * National University of Singapore(新加坡国立大学) Duke-NUS Medical School(杜克-国立新加坡大学医学院) Duke University(杜克大学) Peking University(北京大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 21 pages, 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12051 2025-10-03 cs.CL 57%

TLUE: A Tibetan Language Understanding Evaluation Benchmark

Fan Gao, Cheng Huang, Nyima Tashi, Xiangxiang Wang, Thupten Tsering, Ban Ma-bao, Renzeg Duojie, Gadeng Luosang, Rinchen Dongrub, Dorje Tashi, Hao Wang Xiao Feng, Yongbin Yu

机构 * University of Electronic Science and Technology of China(电子科技大学) Tibet University(西藏大学) University of Texas Southwestern Medical Center(德克萨斯西南医学中心) Southern Methodist University(南方 Methodist 大学) The State Key Laboratory of Tibetan Intelligence(藏语智能国家重点实验室) University of Connecticut(康涅狄格大学)

专题命中 安全评测 :safety(abstract);分类 cs.CL

Comments Accepted by EMNLP Main Conference (Poster)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01030 2025-10-02 cs.AI 57%

Uncovering the Computational Ingredients of Human-Like Representations in LLMs

Zach Studdiford, Timothy T. Rogers, Kushin Mukherjee, Siddharth Suresh

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Stanford University(斯坦福大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00799 2025-10-02 cs.CR cs.AI 57%

Fast, Secure, and High-Capacity Image Watermarking with Autoencoded Text Vectors

Gautier Evennou, Vivien Chappelier, Ewa Kijak

机构 * IRISA, Univ. Rennes, CNRS(IRISA、里昂大学、CNRS) Imatag LABEL4.AI

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00627 2025-10-02 cs.AI 57%

Collaborative-Distilled Diffusion Models (CDDM) for Accelerated and Lightweight Trajectory Prediction

Bingzhang Wang, Kehua Chen, Yinhai Wang

机构 * Department of Civil and Environmental Engineering, University of Washington(华盛顿大学土木与环境工程系)

专题命中 安全评测 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00619 2025-10-02 cs.RO cs.AI 57%

What Did I Learn? Operational Competence Assessment for AI-Based Trajectory Planners

Michiel Braat, Maren Buermann, Marijke van Weperen, Jan-Pieter Paardekooper

机构 * Netherlands Organisation for Applied Scientific Research(荷兰应用科学研究院) Integrated Vehicle Safety Group(集成车辆安全组) Radboud University(拉德堡德大学) Donders Institute for Brain, Cognition and Behaviour(多纳尔斯脑、认知与行为研究所)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted for publication in proceedings of the 2025 IEEE International Automated Vehicle Validation Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00324 2025-10-02 cs.SE cs.IR cs.LG 57%

Which Programming Language and Model Work Best With LLM-as-a-Judge For Code Retrieval?

Lucas Roberts, Denisa Roberts

机构 * Independent Researcher(独立研究者) New York University(纽约大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

Comments Accepted as a full paper at SIGIR-AP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00067 2025-10-02 cs.CV cs.AI cs.HC 57%

Intelligent 5S Audit: Application of Artificial Intelligence for Continuous Improvement in the Automotive Industry

Rafael da Silva Maciel, Lucio Veraldo

机构 * Institute of Science and Technology(科学与技术研究所) Federal University of São Paulo(圣保罗联邦大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 8 pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26106 2025-10-02 cs.RO cs.AI 57%

Autonomous Multi-Robot Infrastructure for AI-Enabled Healthcare Delivery and Diagnostics

Nakhul Kalaivanan, Senthil Arumugam Muthukumaraswamy, Girish Balasubramanian

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 11 pages, 5 figures, MSc dissertation submission draft, prepared for conference/journal consideration

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14520 2025-10-02 cs.AI 57%

What if Othello-Playing Language Models Could See?

Xinyi Chen, Yifei Yuan, Jiaang Li, Serge Belongie, Maarten de Rijke, Anders Søgaard

机构 * University of Amsterdam(阿姆斯特丹大学) ETH Zürich(苏黎世联邦理工学院) University of Copenhagen(哥本哈根大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments ICML 2025 Assessing World Models Workshop; EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26640 2025-10-01 cs.LG cs.CR 57%

SPATA: Systematic Pattern Analysis for Detailed and Transparent Data Cards

João Vitorino, Eva Maia, Isabel Praça, Carlos Soares

机构 * GECAD, ISEP, Polytechnic of Porto Faculty of Engineering, University of Porto(GECAD、ISEP、波尔图理工学院工程学院、波尔图大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

Comments 16 pages, 3 tables, 6 figures, SynDAiTE, ECML PKDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26636 2025-10-01 cs.LG 57%

AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond

Shangding Gu, Xiaohan Wang, Donghao Ying, Haoyu Zhao, Runing Yang, Ming Jin, Boyi Li, Marco Pavone, Serena Yeung-Levy, Jun Wang, Dawn Song, Costas Spanos

机构 * UC Berkeley(伯克利大学) Stanford(斯坦福大学) UCL(伦敦大学学院) Virginia Tech(弗吉尼亚理工学院) Nvidia(英伟达公司)

专题命中 安全评测 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26276 2025-10-01 cs.CL cs.SD 57%

Optimizing Speech Language Models for Acoustic Consistency

Morteza Rohanian, Michael Krauthammer

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26225 2025-10-01 cs.CV cs.AI 57%

An Experimental Study on Generating Plausible Textual Explanations for Video Summarization

Thomas Eleftheriadis, Evlampios Apostolidis, Vasileios Mezaris

机构 * IEEE CBMI 2025

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments IEEE CBMI 2025. This is the authors' accepted version. The final publication is available at https://ieeexplore.ieee.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25651 2025-10-01 cs.AI 57%

AutoLabs: Cognitive Multi-Agent Systems with Self-Correction for Autonomous Chemical Experimentation

Gihan Panapitiya, Emily Saldanha, Heather Job, Olivia Hess

机构 * Pacific Northwest National Laboratory(太平洋西北国家实验室)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25637 2025-10-01 cs.LG 57%

How Does Preconditioning Guide Feature Learning in Deep Neural Networks?

Kotaro Yoshida, Atsushi Nitanda

机构 * Institute of Science Tokyo(东京科学研究院) Agency for Science, Technology and Research (A*STAR)(科技研究局) Nanyang Technological University(南洋理工大学)

专题命中 安全评测 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20411 2025-10-01 cs.CR cs.AI 57%

Adversarial Defense in Cybersecurity: A Systematic Review of GANs for Threat Detection and Mitigation

Tharcisse Ndayipfukamiye, Jianguo Ding, Doreen Sebastian Sarwatt, Adamu Gaston Philipo, Huansheng Ning

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 36 pages, 10 tables, 4figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23326 2025-10-01 cs.HC cs.CY 57%

Designing the Future of Entrepreneurship Education: Exploring an AI-Empowered Scaffold System for Business Plan Development

Junhua Zhu, Lan Luo

专题命中 安全评测 :alignment(abstract);分类 cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.09349 2025-10-01 cs.AI cs.HC cs.MA 57%

Establishing Shared Query Understanding in an Open Multi-Agent System

Nikolaos Kondylidis, Ilaria Tiddi, Annette ten Teije

机构 * Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 9 pages. International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2023), London, United Kingdom

Journal ref In Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems (AAMAS 23). International Foundation for Autonomous Agents and Multiagent Systems, Richland

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24816 2025-09-30 cs.CL 57%

KnowGuard: Knowledge-Driven Abstention for Multi-Round Clinical Reasoning

Xilin Dang, Kexin Chen, Xiaorui Su, Ayush Noori, Iñaki Arango, Lucas Vittor, Xinyi Long, Yuyang Du, Marinka Zitnik, Pheng Ann Heng

专题命中 安全评测 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24443 2025-09-30 cs.AI cs.ET cs.SE 57%

A Systematic Review of Digital Twin-Driven Predictive Maintenance in Industrial Engineering: Taxonomy, Architectural Elements, and Future Research Directions

Leila Ismail, Abdelmoneim Abdelmoti, Arkaprabha Basu, Aymen Dia Eddine Berini, Mohammad Naouss

机构 * Intelligent Distributed Computing and Systems (INDUCE) Lab, Department of Computer Science and Software Engineering, College of Information Technology, United Arab Emirates University, Al-Ain, United Arab Emirates, Emirates Center for Mobility Research, United Arab Emirates University, Al-Ain, United Arab Emirates(智能分布式计算与系统实验室,计算机科学与软件工程系,信息科技学院,阿联酋大学,阿恩,阿联酋,移动性研究中心,阿联酋大学,阿恩,阿联酋) Department of Architectural Engineering, College of Engineering, United Arab Emirates University, Al-Ain, United Arab Emirates(建筑工程系,工程学院,阿联酋大学,阿恩,阿联酋) Department of Information Systems and Security, College of Information Technology, United Arab Emirates University, Al-Ain, United Arab Emirates(信息系统与安全系,信息科技学院,阿联酋大学,阿恩,阿联酋) Department of Computer and Network Engineering, United Arab Emirates University, Al-Ain, United Arab Emirates(计算机与网络工程系,阿联酋大学,阿恩,阿联酋)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24127 2025-09-30 cs.AI cs.DB 57%

Transparent, Evaluable, and Accessible Data Agents: A Proof-of-Concept Framework

Nooshin Bahador

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 20 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22582 2025-09-30 cs.CL 57%

Fine-Grained Detection of Context-Grounded Hallucinations Using LLMs

Yehonatan Peisakhovsky, Zorik Gekhman, Yosi Mass, Liat Ein-Dor, Roi Reichart

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03418 2025-09-30 cs.CL 57%

Which Words Matter Most in Zero-Shot Prompts?

Nikta Gohari Sadr, Sangmitra Madhusudan, Hassan Sajjad, Ali Emami

机构 * Brock University(布罗克大学) Dalhousie University(达尔豪斯大学) Emory University(埃默里大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL

Comments 8 pages (excluding references)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23757 2025-09-30 cs.AI cs.CV 57%

Transparent Visual Reasoning via Object-Centric Agent Collaboration

Benjamin Teoh, Ben Glocker, Francesca Toni, Avinash Kori

机构 * Imperial College London, UK(伦敦帝国学院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22834 2025-09-30 cs.NI cs.AI 57%

Bridging Language Models and Formal Methods for Intent-Driven Optical Network Design

Anis Bekri, Amar Abane, Abdella Battou, Saddek Bensalem

机构 * National Institute of Standards and Technology(美国国家标准技术研究院)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments Accepted at AICCSA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏