The Secret Agenda: LLMs Strategically Lie and Our Current Safety Tools Are Blind
机构 * Independent Researcher(独立研究者)
专题命中 安全评测 :safety(title);分类 cs.AI、cs.CY、cs.LG
Comments 9 pages plus citations and appendix, 7 figures
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Independent Researcher(独立研究者)
专题命中 安全评测 :safety(title);分类 cs.AI、cs.CY、cs.LG
Comments 9 pages plus citations and appendix, 7 figures
机构 * State Key Laboratory of Virtual Reality Technology and Systems, School of Computer Science and Engineering, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,计算机科学与工程学院,北京航空航天大学) ; Hangzhou Innovation Institute of Beihang University, Zhejiang Key Laboratory of Industrial Big Data and Robot Intelligent Systems, Hangzhou, China(北京航空航天大学杭州创新研究院,浙江省工业大数据与机器人智能系统重点实验室) ; Center for AI Business Innovation, Department of Management Science and Systems, University at Buffalo, Buffalo, New York, USA(人工智能商业创新中心,管理科学与系统系,布法罗大学) ; University of North Texas, Denton, Texas, USA(德克萨斯大学达文波特分校) ; Zhongguancun Laboratory, Beijing, China(中关村实验室)
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG
Comments 28 pages, 32 figures, accepted to the Findings of EMNLP 2025
机构 * Synkrasis Labs(Synkrasis实验室) ; Harbin Institute of Technology(哈尔滨工业大学)
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI
Comments Accepted: LAW 2025 Workshop NeurIPS 2025
机构 * University College London(伦敦大学学院)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
Comments 18 pages, 21 figures
机构 * Harbin Institute of Technology(哈尔滨工业大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
机构 * DISI, University of Trento(特伦托大学DISI中心) ; MaiNLP, Center for Information and Language Processing, LMU Munich(慕尼黑大学信息与语言处理中心) ; Munich Center for Machine Learning (MCML), Munich, Germany(慕尼黑机器学习中心) ; Free University of Bozen-Bolzano, Italy(博兹纳自由大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments EMNLP 2025 Main, 38 pages, 33 figures
机构 * Stability AI ; SketchX, University of Surrey(SketchX,大学)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
Comments Project Page: https://hmrishavbandy.github.io/sd35flash/
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling
机构 * Peking University(北京大学) ; LLM-Core
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Comments 11 pages, 2 figures, 2 tables
机构 * Embedded Systems Group (Dept.-E)(嵌入式系统组) ; Virtual Vehicle Research GmbH(虚拟车辆研究公司) ; Control Systems Group (Dept.-E)(控制系统组) ; Institute of Visual Computing(视觉计算研究所) ; Graz University of Technology(格拉茨技术大学)
专题命中 安全评测 :safety(abstract)
Comments Accepted for presentation at the 2025 IEEE International Automated Vehicle Validation Conference (IAVVC 2025). Final version to appear in IEEE Xplore