arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-21 至 2025-08-21 共收录 24 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 2 篇

2504.16133 2025-08-21 cs.CY cs.AI 62%

A Conceptual Framework for AI-based Decision Systems in Critical Infrastructures

Milad Leyli-abadi, Ricardo J. Bessa, Jan Viebahn, Daniel Boos, Clark Borst, Alberto Castagna, Ricardo Chavarriaga, Mohamed Hassouna, Bruno Lemetayer, Giulia Leto, Antoine Marot, Maroua Meddeb, Manuel Meyer, Viola Schiaffonati, Manuel Schneider, Toni Waefler

机构 * IRT SystemX(IRT系统X) INESC TEC TenneT SBB Delft University of Technology(代尔夫特理工大学) EnliteAI Zurich University of Applied Science(苏黎世应用科学大学) Réseau de Transport d’Electricité(电力传输网络) Flatland Association(Flatland协会) Politecnico di Milano(米兰理工大学) University of Applied Sciences Northwestern Switzerland(瑞士西北应用科学大学) Fraunhofer IEE(弗劳恩霍夫研究所)

专题命中 安全训练 :safety(abstract);分类 cs.AI、cs.CY

Journal ref 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14235 2025-08-21 cs.RO 50%

SLAM-based Safe Indoor Exploration Strategy

Omar Mostafa, Nikolaos Evangeliou, Anthony Tzes

机构 * Center for Artificial Intelligence \& Robotics (CAIR) New York University Abu Dhabi (NYUAD) United Arab Emirates Robotics \& Intelligent Systems Control Lab NYUAD United Arab Emirates

专题命中 安全训练 :safety(abstract)

Comments 5 pages, 8 figures. Published in the 2025 11th International Conference on Automation, Robotics, and Applications (ICARA)

Journal ref 2025 11th International Conference on Automation, Robotics, and Applications (ICARA), pp. 375-379

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 越狱攻击 2 篇

2508.14128 2025-08-21 cs.CR cs.AI 85%

CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection

Jiaming Hu, Haoyu Wang, Debarghya Mukherjee, Ioannis Ch. Paschalidis

机构 * Department. of Math & Statistics, Boston University(数学与统计学系,波士顿大学) Department. of Computer Science, University at Albany(计算机科学系,阿尔巴尼大学) Department. of ECE & Systems Eng., Department. of Biomedical Eng., Faculty of Computing & Data Sciences, Boston University(电子工程与系统工程系,生物医学工程系,计算与数据科学学院,波士顿大学)

专题命中 越狱攻击 :jailbreak(title,abstract);safety(abstract);prompt injection(abstract);分类 cs.AI

Comments 11 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14853 2025-08-21 cs.LG 81%

Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent

Sajib Biswas, Mao Nishino, Samuel Jacob Chacko, Xiuwen Liu

机构 * Department of Computer Science, Florida State University(计算机科学系,佛罗里达州立大学) Department of Mathematics, Florida State University(数学系,佛罗里达州立大学)

专题命中 越狱攻击 :alignment(abstract);RLHF(abstract);safety(abstract);jailbreak(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 提示注入 1 篇

2508.14231 2025-08-21 cs.CY cs.AI 73%

Incident Analysis for AI Agents

Carson Ezell, Xavier Roberts-Gaal, Alan Chan

专题命中 提示注入 :safety(abstract);prompt injection(abstract);分类 cs.AI、cs.CY

Comments 16 pages (10 pages main text), 4 figures, 3 tables. To be published in the Proceedings of the 2025 AAAI/ACM Conference on AI, Ethics, & Society (AIES)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 幻觉与事实性 1 篇

2508.14266 2025-08-21 cs.CV cs.AI 57%

Effect of Data Augmentation on Conformal Prediction for Diabetic Retinopathy

Rizwan Ahamed, Annahita Amireskandari, Joel Palko, Carol Laxson, Binod Bhattarai, Prashnna Gyawali

机构 * West Virginia University(西弗吉尼亚大学) University of Aberdeen(阿伯丁大学)

专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI

Comments 3rd Workshop in Data Engineering in Medical Imaging (DEMI), MICCAI-2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 隐私与版权 1 篇

2502.04519 2025-08-21 eess.AS cs.LG 57%

GenVC: Self-Supervised Zero-Shot Voice Conversion

Zexin Cai, Henry Li Xinyuan, Ashi Garg, Leibny Paola García-Perera, Kevin Duh, Sanjeev Khudanpur, Matthew Wiesner, Nicholas Andrews

机构 * Human Language Technology Center of Excellence Johns Hopkins University(人类语言技术中心杰出成就约翰霍普金斯大学)

专题命中 隐私与版权 :alignment(abstract);分类 cs.LG

Comments accepted by 2025 IEEE ASRU

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 安全评测 9 篇

2508.14735 2025-08-21 cs.CL cs.AI 81%

Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference

Samir Abdaljalil, Erchin Serpedin, Khalid Qaraqe, Hasan Kurban

机构 * Texas A\&M University, College Station, TX., USA(德克萨斯大学) Hamad Bin Khalifa University, Doha, Qatar(哈马德·本·卡伊夫大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14741 2025-08-21 cs.LG 79%

CaTE Data Curation for Trustworthy AI

Mary Versa Clemens-Sewall, Christopher Cervantes, Emma Rafkin, J. Neil Otte, Tom Magelinski, Libby Lewis, Michelle Liu, Dana Udwin, Monique Kirkman-Bey

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19422 2025-08-21 cs.CY cs.AI cs.HC 62%

Generative AI in K-12 Education: The CyberScholar Initiative

Vania Castro, Ana Karina de Oliveira Nascimento, Raigul Zheldibayeva, Duane Searsmith, Akash Saini, Bill Cope, Mary Kalantzis

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14654 2025-08-21 cs.AI 57%

Entropy-Constrained Strategy Optimization in Urban Floods: A Multi-Agent Framework with LLM and Knowledge Graph Integration

Peilin Ji, Xiao Xue, Simeng Wang, Wenhao Yan

机构 * College of Intelligence and Computing(智能与计算学院)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 17 pages including appendix, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14427 2025-08-21 cs.CL 57%

Knowledge Graph-Infused Fine-Tuning for Structured Reasoning in Large Language Models

Wuyang Zhang, Yexin Tian, Xiandong Meng, Mengjie Wang, Junliang Du

专题命中 安全评测 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11786 2025-08-21 stat.ML cs.LG 57%

Parallelly Tempered Generative Adversarial Nets: Toward Stabilized Gradients

Jinwon Sohn, Qifan Song

机构 * Booth School of Business, University of Chicago(芝加哥大学商学院) Department of Statistics, Purdue University(普渡大学统计学系)

专题命中 安全评测 :trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21755 2025-08-21 cs.CV 50%

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

Dian Zheng, Ziqi Huang, Hongbo Liu, Kai Zou, Yinan He, Fan Zhang, Lulu Gu, Yuanhan Zhang, Jingwen He, Wei-Shi Zheng, Yu Qiao, Ziwei Liu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) S-Lab, Nanyang Technological University(南洋理工大学S实验室) Sun Yat-sen University(中山大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 安全评测 :alignment(abstract)

Comments Equal contributions from first two authors. Project page: https://vchitect.github.io/VBench-2.0-project/ Code: https://github.com/Vchitect/VBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04171 2025-08-21 cs.CV 50%

DuCos: Duality Constrained Depth Super-Resolution via Foundation Model

Zhiqiang Yan, Zhengxue Wang, Haoye Dong, Jun Li, Jian Yang, Gim Hee Lee

机构 * National University of Singapore(国立新加坡大学) Nanjing University of Science and Technology(南京理工大学)

专题命中 安全评测 :alignment(abstract)

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02028 2025-08-21 cs.CV 50%

Bench2ADVLM: A Closed-Loop Benchmark for Vision-language Models in Autonomous Driving

Tianyuan Zhang, Ting Jin, Lu Wang, Jiangfan Liu, Siyuan Liang, Mingchuan Zhang, Aishan Liu, Xianglong Liu

机构 * Beihang University(北航大学) Nanyang Technological University(南洋理工大学) Henan University of Science and Technology(河南科技大学)

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏

7. AI治理与伦理 2 篇

2508.14116 2025-08-21 cs.CY cs.AI 62%

Enriching Moral Perspectives on AI: Concepts of Trust amongst Africans

Lameck Mbangula Amugongo, Nicola J Bidwell, Joseph Mwatukange

机构 * Namibia University of Science \& Technology 13 Jackson Kaujeua Windhoek Namibia 9000 Rhodes University Makhanda South Africa International University of Management Namibia Charles Darwin University Australia Namibia University of Science \& Technology Rhodes University International University of Management Charles Darwin University

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14415 2025-08-21 cs.AI 57%

The Agent Behavior: Model, Governance and Challenges in the AI Digital Age

Qiang Zhang, Pei Yan, Yijia Xu, Chuanpo Fu, Yong Fang, Yang Liu

机构 * School of Cyber Science and Engineering, Sichuan University, China(计算机科学与工程学院,四川大学) College of Computing and Data Science, Nanyang Technological University, Sinapore(计算与数据科学学院,南洋理工大学) College of Electronics and Information Engineering, Shenzhen University, China(电子与信息工程学院,深圳大学) Department of Computer Science and Technology, Tsinghua University, China(计算机科学与技术系,清华大学)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏

8. 其他安全 6 篇

2508.14618 2025-08-21 cs.LG 57%

A Fuzzy-Enhanced Explainable AI Framework for Flight Continuous Descent Operations Classification

Amin Noroozi, Sandaruwan K. Sethunge, Elham Norouzi, Phat T. Phan, Kavinda U. Waduge, Md. Arafatur Rahman

专题命中 其他安全 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14214 2025-08-21 cs.AI 57%

Large Language Models are Highly Aligned with Human Ratings of Emotional Stimuli

Mattson Ogg, Chace Ashcraft, Ritwik Bose, Raphael Norman-Tenazas, Michael Wolmetz

专题命中 其他安全 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14729 2025-08-21 cs.CV 50%

Multiscale Video Transformers for Class Agnostic Segmentation in Autonomous Driving

Leila Cheshmi, Mennatullah Siam

机构 * Ontariotechu(安大略理工学院)

专题命中 其他安全 :safety(abstract)

Comments 6 pages, 2 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03984 2025-08-21 cs.CV 50%

CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning

Jeonghyo Song, Kimin Yun, DaeUng Jo, Jinyoung Kim, Youngjoon Yoo

机构 * Department of Artificial Intelligence, Chung-Ang University(Chung-Ang大学人工智能系) Visual Intelligence Lab., ETRI(ETRI视觉智能实验室) University of Science and Technology (UST)(科技大学(UST)) School of Electronics Engineering, Kyungpook National University(Kyungpook国立大学电子工程学院)

专题命中 其他安全 :safety(abstract)

Comments 6 pages, 3 figures. Accepted at IEEE International Conference on Advanced Visual and Signal-Based Systems 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00319 2025-08-21 eess.SY cs.SY 50%

Beyond Quadratic Costs: A Bregman Divergence Approach to H$_\infty$ Control

Joudi Hajar, Reza Ghane, Babak Hassibi

专题命中 其他安全 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.07563 2025-08-21 cs.IR 50%

Reinforcement Learning to Rank Using Coarse-grained Rewards

Yiteng Tu, Zhichao Xu, Tao Yang, Weihang Su, Yujia Zhou, Yiqun Liu, Fen Lin, Qin Liu, Qingyao Ai

专题命中 其他安全 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏