Multivariable Fractional Polynomials for lithium-ion batteries degradation models under dynamic conditions
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
Comments 45 pages, 7 figures, 3 tables
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
Comments 45 pages, 7 figures, 3 tables
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
Journal ref Secure and Private Systems for Machine Learning Workshop 2021
专题命中 越狱攻击 :trustworthy(abstract);分类 cs.CL
Comments EMNLP-BlackboxNLP, 2021
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
专题命中 越狱攻击 :safety(abstract);分类 cs.CY
Comments 3 pages, 6 figures
Journal ref International Journal of Engineering and Advanced Technology, Volume-8 Issue-4S2, April 2019
专题命中 越狱攻击 :safety(abstract);分类 cs.AI
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
Comments Accepted at AAAI'20
Journal ref AAAI 2020
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
Comments 14 pages, 4 tables, 3 figures, 5 equations, In Proceeding of WISA 2020 (THE 21ST WORLD CONFERENCE ON INFORMATION SECURITY APPLICATIONS)
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
Comments Accepted by CVPR 2020
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
Comments survey, adversarial attacks, defenses
专题命中 越狱攻击 :alignment(abstract);分类 cs.LG
Comments 14 figures, 17 pages, Draft
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
Comments CVPR SAIAD - Workshop 2019
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
Journal ref Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pp. 52-68, 2018
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
Comments 10 pages, 2 figures, 4 tables
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
Comments Project page: http://yclin.me/RL_attack_detection/ Code: https://github.com/yenchenlin/rl-attack-detection
专题命中 越狱攻击 :分类 cs.CL、cs.AI、cs.LG;DPO(comments)
Comments ICML 2024. Code available at https://github.com/ruocwang/dpo-diffusion
Journal ref Proceedings of the 41st International Conference on Machine Learning (ICML 2024)
$PC^2$:通过基于GPT的文本到图像模型的越狱攻击生成政治争议内容
专题命中 越狱攻击 :safety(abstract)
AI总结 提出首个黑盒政治越狱框架$PC^2$,利用身份保留描述映射和地缘政治远距离翻译绕过安全过滤器,在GPT系列模型上实现高达86%的攻击成功率。
Comments To appear in the 33rd ACM Conference on Computer and Communications Security (CCS 2026)
为视觉几何接地Transformer生成多视角对抗样本
机构 * Hong Kong Baptist University(香港浸会大学)
专题命中 越狱攻击 :alignment(abstract)
AI总结 本研究针对3D基础模型VGGT的安全漏洞,提出MVAP-G多视角对抗扰动生成器,通过跨视角对抗对齐机制生成一致扰动,可无需迭代优化即可显著降低VGGT性能,为3D视觉系统鲁棒性研究提供了新方向。
Comments ECCV 2026
网联车辆中的自主网络防御:一种面向V2X安全的多智能体方法
专题命中 越狱攻击 :safety(abstract)
AI总结 针对网联车辆V2X安全,提出三层多智能体架构,以SAE标准时序为硬性约束,分层处理消息分类、冲突解决与模型优化,解决现有入侵检测系统的安全-安保冲突问题。
PROVE:基于可验证证据的无训练提示词恢复
专题命中 越狱攻击 :alignment(abstract)
AI总结 该研究提出无训练的黑盒提示词反转攻击PROVE,通过可验证场景描述重建提示词,在多数据集上优于基线方法,可用于版权保护相关研究。
SoK:代理AI的攻击面——工具与自主性
专题命中 越狱攻击 :prompt injection(abstract)
AI总结 本文系统分析了代理AI的安全风险,提出攻击分类与评估指标,探讨防御措施及开放研究挑战,帮助研究人员和工程师应对代理AI的安全威胁。
超越输入防护:为执行感知攻击检测重建跨代理语义流
专题命中 越狱攻击 :prompt injection(abstract)
AI总结 SysName通过重建跨代理语义流,实现执行感知的攻击检测,有效识别多种复合攻击向量,提升多代理系统安全性。
Comments 16 pages, 15 figures
针对视频大语言模型中提示引导采样的投毒攻击
机构 * National University of Singapore(新加坡国立大学) ; University of New South Wales(新南威尔士大学) ; CSIRO’s Data61(CSIRO数据61)
专题命中 越狱攻击 :safety(abstract)
AI总结 该研究针对视频大语言模型的提示引导采样提出PoisonVID投毒攻击,在多种模型与采样器组合上实现高攻击成功率,揭示了PGS存在的结构性安全隐患。
Comments 16 pages, 5 figures
SoK:面向意图的多轮大语言模型越狱系统化研究
专题命中 越狱攻击 :safety(abstract)
AI总结 本研究提出面向意图的多轮LLM越狱分类法,发现攻击有效性取决于意图组织精细度,推动检测范围升级,表明轮级安全机制不足,需对应层级评估协议。
从角色提示到无限思考:利用角色条件进行大语言模型推理成本攻击
专题命中 越狱攻击 :alignment(abstract)
AI总结 研究LLMs推理成本问题,提出RolePlay框架利用角色条件诱导低效行为放大推理成本,实验证明该框架优于现有方法,为LLM推理效率攻击提供新视角。
Comments 17pages
检测器学习到错误的东西:针对物理可实现攻击的抗捷径对抗训练
机构 * School of Transportation Science and Engineering, Beihang University(北京航空航天大学交通科学与工程学院) ; State Key Lab of Intelligent Transportation System(智能交通系统国家重点实验室) ; Zhongguancun Laboratory(中关村实验室) ; School of Systems Science and Engineering, Sun Yat-Sen University(中山大学系统科学与工程学院) ; Shandong Sino-Aisa Tire Proving Ground Co.,Ltd.(山东中亚洲轮胎试验场有限公司)
专题命中 越狱攻击 :safety(abstract)
AI总结 研究针对物理可实现攻击,提出实例级对比对抗训练框架InsCAT,通过SICA、ROPO和Guard等方法防止检测器将对抗纹理用作独立决策线索,经实验验证其在多场景和检测器上效果良好,提升检测可靠性。
当HTTP 402遇上区块链:新兴x402支付的风险
专题命中 越狱攻击 :safety(abstract)
AI总结 研究x402支付协议,通过系统研究确定促进者的安全规则,基于规则违规分析得出新攻击向量,提出半自动黑盒工具评估安全性,发现部署均有违规,还通过实证测量量化相关指标,揭示其安全风险。
Journal ref USENIX Security 2026
VENOMREC: 多模态交互性污染用于多模态大语言模型推荐系统中的定向推广
专题命中 越狱攻击 :alignment(abstract)
AI总结 VENOMREC通过跨模态交互性污染技术,在多模态大语言模型推荐系统中实现定向推广,有效提升推荐性能。