arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2025-08-20 至 2025-08-20 共收录 2 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 越狱攻击 2 篇

2508.09288 2025-08-20 cs.CR cs.AI cs.CL 73%

Can AI Keep a Secret? Contextual Integrity Verification: A Provable Security Architecture for LLMs

Aayush Gupta

机构 * Aayush Gupta(独立研究者)

专题命中 越狱攻击 :jailbreak(abstract);prompt injection(abstract);分类 cs.CL、cs.AI

Comments 2 figures, 3 tables; code and certification harness: https://github.com/ayushgupta4897/Contextual-Integrity-Verification ; Elite-Attack dataset: https://huggingface.co/datasets/zyushg/elite-attack

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02332 2025-08-20 cs.CR 50%

PII Jailbreaking in LLMs via Activation Steering Reveals Personal Information Leakage

Krishna Kanth Nakka, Xue Jiang, Dmitrii Usynin, Xuebing Zhou

专题命中 越狱攻击 :alignment(abstract)

Comments Preprint. V2 Updated with dataset filtering, benchmarking privacy evaluator and additional latent space visualizations

详情

展开后加载摘要…

URL PDF HTML 收藏