Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
机构 * Rutgers University(新泽西罗格斯大学)
专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Rutgers University(新泽西罗格斯大学)
专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CL、cs.LG
机构 * Rochester Institute of Technology(罗切斯特技术研究所)
专题命中 安全评测 :alignment(title,abstract);分类 cs.LG
机构 * York University(约克大学) ; Vector Institute for AI(人工智能矢量研究所) ; Dialpad Inc.(Dialpad公司) ; Nanyang Technological University(南洋理工大学) ; Salesforce AI Research(Salesforce人工智能研究)
专题命中 安全评测 :alignment(abstract);safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI
Comments Survey paper; 45 data science agents; under review
机构 * University of Chicago(芝加哥大学)
专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.CY
Comments Our code is available, see https://github.com/ChicagoHAI/self-recognition
专题命中 安全评测 :trustworthy(title);分类 cs.LG
Comments 1tables,6 figs,11pages
机构 * Department of Statistics, Eskisehir Technical University, Turkiye(埃斯基谢普大学统计系) ; Leiden Institute of Advanced Computer Science, Leiden University, the Netherlands(莱顿大学高级计算机科学研究所) ; Faculty of Mathematics and Information Science, Warsaw University of Technology, Poland(华沙理工大学数学与信息科学学院) ; Informatics and Mechanics, University of Warsaw, Faculty of Mathematics, Poland(华沙大学信息技术与力学系)
专题命中 安全评测 :trustworthy(title);分类 cs.LG
Comments Accepted at 28th International Conference on Discovery Science 2025
Journal ref In: Džeroski, S., Levatić, J., Pio, G., Simidjievski, N. (eds) Discovery Science. DS 2025. Lecture Notes in Computer Science, vol 16090. Springer, Cham
机构 * Institute for Infocomm Research, Agency for Science, Technology and Research, Singapore(新加坡资讯通信研究院,科技研究局) ; Propulsion and Space Research Center, Technology Innovation Institute, UAE(阿联酋技术创新研究所推进与航天研究中心) ; James Watt School of Engineering, University of Glasgow, UK(格拉斯哥大学詹姆斯·瓦特工程学院) ; ISTD Pillar at SUTD(新加坡科技设计大学ISTD支柱)
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG
机构 * Arizona State University(亚利桑那州立大学) ; University of Kansas(堪萨斯大学) ; University of Notre Dame(圣母大学) ; Duke University(杜克大学)
专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.AI
机构 * Tsinghua University(清华大学) ; The Ohio State University(俄亥俄州立大学) ; UC Berkeley(加州大学伯克利分校)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments Published in ICLR 2024
机构 * Department of Data and Decision Science, Technion - Israel Institute of Technology(数据与决策科学系,技术离子理工学院) ; Department of Computer Science, Technion - Israel Institute of Technology(计算机科学系,技术离子理工学院)
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.LG
Comments Accepted to NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle: Benchmarks, Emergent Abilities, and Scaling
机构 * School of Physical and Mathematical Sciences, Nanyang Technological University(南洋理工大学物理与数学科学学院) ; Lab of Artificial Intelligence for Education, East China Normal University(华东师范大学教育人工智能实验室) ; School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) ; State Key Laboratory of Robotics and Systems, Harbin Institute of Technology(哈尔滨工业大学机器人系统国家重点实验室) ; College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CY
Comments Preprint, Under review
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Comments 5 figures
Journal ref https://www.mdpi.com/2075-4418/15/19/2486
机构 * Faculty of Information Technology, University of Jyväskylä(信息技术学院,约赫斯库利亚大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Comments Accepted at SEAA 2025. Appearing in Springer LNCS 16081, pages 164-180
Journal ref In: SEAA 2025 proceedings, LNCS vol. 16081, Springer
机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) ; Georgia Institute of Technology(佐治亚理工学院) ; University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) ; TwelveLabs
专题命中 安全评测 :alignment(abstract);分类 cs.AI
机构 * OpenDataLab, Shanghai Artificial Intelligence Laboratory(OpenDataLab,上海人工智能实验室) ; Shanghai Jiao Tong University(上海交通大学) ; Soochow University(苏州大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL
Comments Accepted by NeurIPS2025
专题命中 安全评测 :alignment(abstract);分类 cs.CY
Comments 2025 EMNLP HCI+NLP Workshop Short Paper
机构 * Université de Montréal(蒙特利尔大学) ; Mila – Quebec AI Institute(魁北克人工智能研究所)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
机构 * The University of Hong Kong(香港大学) ; Kling Team, Kuaishou Technology(快手科技 Kling 团队) ; Lingnan University(岭大)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
Comments ICCV 2025
机构 * EverEx ; University of Michigan(密歇根大学) ; ETH Zurich(苏黎世联邦理工学院) ; Yonsei University(延世大学)
专题命中 安全评测 :alignment(abstract)
Comments 18 pages, 7 figures, Project Page:https://hyelinnam.github.io/Cameo/
机构 * Carnegie Mellon University(卡内基梅隆大学) ; Harvard University(哈佛大学) ; University of Alabama at Birmingham(阿拉巴马大学伯明翰分校) ; Oak Ridge National Laboratory(橡树岭国家实验室) ; K. J. Somaiya College of Engineering(K.J. Somaiya 工程学院)
专题命中 安全评测 :alignment(abstract)
机构 * Beijing Normal–Hong Kong Baptist University(北京师范大学-香港 Baptist大学)
专题命中 安全评测 :alignment(abstract)
专题命中 安全评测 :safety(abstract)