arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-09-26 至 2025-09-26 共收录 12 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 12 篇

2509.20393 2025-09-26 cs.CY cs.AI cs.LG 78%

The Secret Agenda: LLMs Strategically Lie and Our Current Safety Tools Are Blind

Caleb DeLeeuw, Gaurav Chawla, Aniket Sharma, Vanessa Dietze

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :safety(title);分类 cs.AI、cs.CY、cs.LG

Comments 9 pages plus citations and appendix, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20680 2025-09-26 cs.LG cs.CL cs.CR 73%

Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation

Wenkai Guo, Xuefeng Liu, Haolin Wang, Jianwei Niu, Shaojie Tang, Jing Yuan

机构 * State Key Laboratory of Virtual Reality Technology and Systems, School of Computer Science and Engineering, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,计算机科学与工程学院,北京航空航天大学) Hangzhou Innovation Institute of Beihang University, Zhejiang Key Laboratory of Industrial Big Data and Robot Intelligent Systems, Hangzhou, China(北京航空航天大学杭州创新研究院,浙江省工业大数据与机器人智能系统重点实验室) Center for AI Business Innovation, Department of Management Science and Systems, University at Buffalo, Buffalo, New York, USA(人工智能商业创新中心,管理科学与系统系,布法罗大学) University of North Texas, Denton, Texas, USA(德克萨斯大学达文波特分校) Zhongguancun Laboratory, Beijing, China(中关村实验室)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

Comments 28 pages, 32 figures, accepted to the Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20998 2025-09-26 cs.AI 70%

CORE: Full-Path Evaluation of LLM Agents Beyond Final State

Panagiotis Michelakis, Yiannis Hadjiyiannis, Dimitrios Stamoulis

机构 * Synkrasis Labs(Synkrasis实验室) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

Comments Accepted: LAW 2025 Workshop NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21287 2025-09-26 cs.CL cs.AI 62%

DisCoCLIP: A Distributional Compositional Tensor Network Encoder for Vision-Language Understanding

Kin Ian Lo, Hala Hawashin, Mina Abbaszadeh, Tilen Limback-Stokin, Hadi Wazni, Mehrnoosh Sadrzadeh

机构 * University College London(伦敦大学学院)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20520 2025-09-26 cs.AI cs.DC cs.LG 62%

Adaptive Approach to Enhance Machine Learning Scheduling Algorithms During Runtime Using Reinforcement Learning in Metascheduling Applications

Samer Alshaer, Ala Khalifeh, Roman Obermaisser

专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG

Comments 18 pages, 21 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20378 2025-09-26 cs.CL cs.AI 62%

Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation

Sirui Wang, Andong Chen, Tiejun Zhao

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11771 2025-09-26 cs.CL cs.AI 62%

The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It

Leonardo Bertolazzi, Philipp Mondorf, Barbara Plank, Raffaella Bernardi

机构 * DISI, University of Trento(特伦托大学DISI中心) MaiNLP, Center for Information and Language Processing, LMU Munich(慕尼黑大学信息与语言处理中心) Munich Center for Machine Learning (MCML), Munich, Germany(慕尼黑机器学习中心) Free University of Bozen-Bolzano, Italy(博兹纳自由大学)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 Main, 38 pages, 33 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21318 2025-09-26 cs.CV cs.AI 57%

SD3.5-Flash: Distribution-Guided Distillation of Generative Flows

Hmrishav Bandyopadhyay, Rahim Entezari, Jim Scott, Reshinth Adithyan, Yi-Zhe Song, Varun Jampani

机构 * Stability AI SketchX, University of Surrey(SketchX,大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments Project Page: https://hmrishavbandy.github.io/sd35flash/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21310 2025-09-26 cs.AI 57%

SAGE: A Realistic Benchmark for Semantic Understanding

Samarth Goel, Reagan J. Lee, Kannan Ramchandran

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21208 2025-09-26 cs.CL 57%

CLaw: Benchmarking Chinese Legal Knowledge in Large Language Models - A Fine-grained Corpus and Reasoning Analysis

Xinzhe Xu, Liang Zhao, Hongshen Xu, Chen Chen

机构 * Peking University(北京大学) LLM-Core

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20418 2025-09-26 cs.CR cs.AI cs.ET 57%

A Taxonomy of Data Risks in AI and Quantum Computing (QAI) - A Systematic Review

Grace Billiris, Asif Gill, Madhushi Bandara

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

Comments 11 pages, 2 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19096 2025-09-26 cs.CV cs.SE 50%

Investigating Traffic Accident Detection Using Multimodal Large Language Models

Ilhan Skender, Kailin Tong, Selim Solmaz, Daniel Watzenig

机构 * Embedded Systems Group (Dept.-E)(嵌入式系统组) Virtual Vehicle Research GmbH(虚拟车辆研究公司) Control Systems Group (Dept.-E)(控制系统组) Institute of Visual Computing(视觉计算研究所) Graz University of Technology(格拉茨技术大学)

专题命中 安全评测 :safety(abstract)

Comments Accepted for presentation at the 2025 IEEE International Automated Vehicle Validation Conference (IAVVC 2025). Final version to appear in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏