arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-15 至 2025-08-15 共收录 13 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 13 篇

2508.09937 2025-08-15 cs.CL cs.AI cs.LG 87%

A Comprehensive Evaluation framework of Alignment Techniques for LLMs

Muneeza Azmat, Momin Abbas, Maysa Malfiza Garcia de Macedo, Marcelo Carpinette Grave, Luan Soares de Souza, Tiago Machado, Rogerio A de Paula, Raya Horesh, Yixin Chen, Heloisa Caroline de Souza Pereira Candello, Rebecka Nordenlow, Aminat Adebiyi

机构 * IBM Research(IBM研究院)

专题命中 安全评测 :alignment(title,abstract);RLHF(abstract);safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments In submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00399 2025-08-15 cs.CV 78%

iSafetyBench: A video-language benchmark for safety in industrial environment

Raiyaan Abdullah, Yogesh Singh Rawat, Shruti Vyas

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 安全评测 :safety(title,abstract)

Comments Accepted to VISION'25 - ICCV 2025 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10146 2025-08-15 cs.AI 70%

Agentic AI Frameworks: Architectures, Protocols, and Design Challenges

Hana Derouiche, Zaki Brahmi, Haithem Mazeni

机构 * University of Kairouan(卡鲁安大学) SMART Lab, University of Tunis(图纳大学SMART实验室) University of Sousse(索斯大学) Riadi Lab, Compus Manouba(曼努巴大学里亚迪实验室) University of Jandouba(贾杜巴大学)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10028 2025-08-15 cs.CL cs.AI cs.HC cs.LG 67%

PREF: Reference-Free Evaluation of Personalised Text Generation in LLMs

Xiao Fu, Hossein A. Rahmani, Bin Wu, Jerome Ramos, Emine Yilmaz, Aldo Lipani

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10494 2025-08-15 cs.LG cs.AI cs.MA 62%

A Unified Multi-Agent Framework for Universal Multimodal Understanding and Generation

Jiulin Li, Ping Huang, Yexin Li, Shuo Chen, Juewen Hu, Ye Tian

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments 8 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10433 2025-08-15 cs.AI cs.CV cs.LG 62%

We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning

Runqi Qiao, Qiuna Tan, Peiqing Yang, Yanzi Wang, Xiaowan Wang, Enhui Wan, Sitong Zhou, Guanting Dong, Yuchen Zeng, Yida Xu, Jie Wang, Chong Sun, Chen Li, Honggang Zhang

机构 * BUPT(北京邮电大学) WeChat Vision, Tencent Inc.(腾讯公司) Tsinghua University(清华大学)

专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10022 2025-08-15 cs.CL cs.AI 62%

Conformal P-Value in Multiple-Choice Question Answering Tasks with Provable Risk Control

Yuanchang Ye

专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13871 2025-08-15 cs.LG cs.AI cs.CR 62%

An Explainable Transformer-based Model for Phishing Email Detection: A Large Language Model Approach

Mohammad Amaz Uddin, Md Mahiuddin, Iqbal H. Sarker

机构 * Department of Computer Science and Engineering, BGC Trust University Bangladesh(Bangladesh BGC Trust 大学 计算机科学与工程系) Department of Computer Science and Engineering, International Islamic University Chittagong(Bangladesh 国际伊斯兰大学 昌德加荣分校 计算机科学与工程系) Centre for Securing Digital Futures, School of Science, Edith Cowan University(澳大利亚 埃德温·考文大学 科学学院 安全数字未来中心)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10869 2025-08-15 cs.CV cs.AI 57%

Medico 2025: Visual Question Answering for Gastrointestinal Imaging

Sushant Gautam, Vajira Thambawita, Michael Riegler, Pål Halvorsen, Steven Hicks

机构 * SimulaMet - Simula Metropolitan Center for Digital Engineering, Oslo, Norway(SimulaMet - Simula Metropolitan Center for Digital Engineering,挪威奥斯陆) Simula Research Laboratory, Oslo, Norway(Simula研究实验室,挪威奥斯陆) OsloMet - Oslo Metropolitan University, Oslo, Norway(OsloMet - 奥斯陆 Metropolitan 大学,挪威奥斯陆)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13307 2025-08-15 cs.CV cs.AI 57%

Quantitative Comparison of Fine-Tuning Techniques for Pretrained Latent Diffusion Models in the Generation of Unseen SAR Images

Solène Debuysère, Nicolas Trouvé, Nathan Letheule, Olivier Lévêque, Elise Colin

机构 * Paris-Saclay University(巴黎-萨克雷大学) ONERA - The French Aerospace Lab(法国航空航天实验室)

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10397 2025-08-15 cs.CV cs.AI 57%

PQ-DAF: Pose-driven Quality-controlled Data Augmentation for Data-scarce Driver Distraction Detection

Haibin Sun, Xinghui Song

机构 * College of Computer Science and Engineering, Shandong University of Science and Technology(计算机科学与工程学院,山东科技大学)

专题命中 安全评测 :safety(abstract);分类 cs.AI

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10358 2025-08-15 cs.AI 57%

What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles

Mengtao Zhou, Sifan Wu, Huan Zhang, Qi Sima, Bang Liu

专题命中 安全评测 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10309 2025-08-15 cs.CV 50%

From Pixel to Mask: A Survey of Out-of-Distribution Segmentation

Wenjie Zhao, Jia Li, Yunhui Guo

机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 安全评测 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏