Red-Teaming Segment Anything Model
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments CVPR 2024 - The 4th Workshop of Adversarial Machine Learning on Computer Vision: Robustness of Foundation Models
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments CVPR 2024 - The 4th Workshop of Adversarial Machine Learning on Computer Vision: Robustness of Foundation Models
专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :prompt injection(abstract);分类 cs.CL、cs.AI
Comments 34 pages, 8 figures Codebase: https://github.com/PromptLabs/hackaprompt Dataset: https://huggingface.co/datasets/hackaprompt/hackaprompt-dataset/blob/main/README.md Playground: https://huggingface.co/spaces/hackaprompt/playground
专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.LG
Comments 32 pages. Implementation available at https://github.com/JonasGeiping/carving
专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.LG
专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments USENIX Security 2024 (https://www.usenix.org/conference/usenixsecurity24/presentation/dahiya)
专题命中 越狱攻击 :safety(abstract);分类 cs.CL、cs.AI
专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.LG
专题命中 越狱攻击 :alignment(abstract);分类 cs.AI、cs.LG
Journal ref Conference on Neural Information Processing Systems (NeurIPS), Neural Information Processing Systems Foundation, Dec 2023, New Orleans (Louisiana), United States
专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.LG
专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL、cs.CY
Comments 13 pages, 9 figures, 7 tables, accepted to findings of EMNLP 2023
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments Technical report
专题命中 越狱攻击 :alignment(abstract);分类 cs.CL、cs.LG
Comments 12 pages
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments Published as a conference paper at the International Conference on Learning Representations (ICLR 2022). Code is available at https://sparseevoattack.github.io/
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.CY
Journal ref Published in ACM CCS 2022. Please cite the CCS version
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments Submitted to conference
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments Accepted at IEEE/CVF WACV 2022 MAP
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments 18 pages, 6 figures
专题命中 越狱攻击 :safety(abstract);分类 cs.AI、cs.LG
Comments Preprint
用于激光雷达距离图像合成的对抗引导扩散
机构 * School of Electrical and Computer Engineering, National Technical University of Athens(雅典国立技术大学电气与计算机工程学院) ; Industrial Systems Institute, Athena Research Center(雅典娜研究中心工业系统研究所)
专题命中 越狱攻击 :safety(abstract);分类 cs.LG;trustworthy(comments)
AI总结 研究针对二维距离图像分割的无限制对抗攻击,提出基于扩散并利用分割损失对抗引导的方法,在SemanticKITTI数据集实验,能跨架构转移,相比基线在有效性与现实性间有独特权衡,实现可控退化。
Comments Accepted at the 1st Workshop on Secure and Trustworthy AI (STAI 2026), co-located with the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD 2026)
机构 * Independent Researcher(独立研究者)
专题命中 越狱攻击 :prompt injection(abstract,comments);分类 cs.AI
Comments 33 pages, 3 figures, 6 tables. Keywords: LLM security; defense-in-depth; prompt injection; activation steering; multimodal sandbox; threat modeling
TEE-X:面向边缘端大视觉模型的感知可信执行环境加速框架
专题命中 越狱攻击 :safety(abstract);分类 cs.AI
AI总结 该研究提出TEE-X框架,在Arm TrustZone的OP-TEE上验证,可在NVIDIA Jetson AGX Xavier上实现大视觉模型的安全高效边缘推理,兼顾性能与安全性。
激活引导用于不牺牲连贯性的对齐开放式生成
机构 * Tara Research ; Technical University of Munich(慕尼黑工业大学) ; Mila Quebec AI Institute(Mila魁北克人工智能研究所)
专题命中 越狱攻击 :alignment(abstract);分类 cs.AI
AI总结 本文提出激活引导方法,通过在生成过程中持续修正偏差激活,提升大模型的对齐性能,实验表明StTP和StMP在保持连贯性的同时更有效维护通用能力。