arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-07-30 至 2025-07-30 共收录 5 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 5 篇

2507.21132 2025-07-30 cs.AI cs.CY cs.LG 75%

Can You Trust an LLM with Your Life-Changing Decision? An Investigation into AI High-Stakes Responses

Joshua Adrian Cahyono, Saran Subramanian

专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21133 2025-07-30 cs.CR cs.AI 70%

Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities

Atil Samancioglu

机构 * Atil Samancioglu(独立研究者)

专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18172 2025-07-30 cs.CR cs.LG 57%

GenAI Security: Outsmarting the Bots with a Proactive Testing Framework

Sunil Kumar Jang Bahadur, Gopala Dhar, Lavi Nigam

机构 * AI \& GenAI Specialist Cloud GTM Google Mumbai, India AI Engineer, AI Services Google Cloud Consulting (GCC) Google Mumbai, India Industry Solutions Google Gurugram, India

专题命中 安全训练 :prompt injection(abstract);分类 cs.LG

Comments IEEE CAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21619 2025-07-30 cs.CV 50%

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO

Wei Guan, Jun Lan, Jian Cao, Hao Tan, Huijia Zhu, Weiqiang Wang

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21547 2025-07-30 math.OC cs.RO cs.SY eess.SY 50%

Decentralized Modeling of Vehicular Maneuvers and Interactions at Urban Junctions

Saeed Rahmani, Simeon C. Calvert, Bart van Arem

机构 * Delft University of Technology(代尔夫特理工大学)

专题命中 安全训练 :safety(abstract)

Comments Manuscript under review

详情

展开后加载摘要…

URL PDF HTML 收藏