SAGA: A Security Architecture for Governing AI Agentic Systems
机构 * Northeastern University(东北大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Northeastern University(东北大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
机构 * AgroParisTech - MIA(阿格罗巴黎技术学院-信息分析中心)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments in French language
机构 * Georgia Institute of Technology(佐治亚理工学院) ; Adobe Research(Adobe研究)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
Comments Accepted to UIST 2025. 18 pages, 9 figures, 2 tables. For a demo video, see https://youtu.be/uobhmxo6EIE
机构 * Zhejiang University(浙江大学) ; ZJU-Angelalign R&D Center for Intelligence Healthcare(浙江大学智能医疗研发中心)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
机构 * Institute for Clarity in Documentation(文档清晰研究所) ; Inria Paris-Rocquencourt(巴黎-罗克琴特研究所) ; Rajiv Gandhi University(拉吉夫·甘地大学) ; Tsinghua University(清华大学) ; Palmer Research Laboratories(帕尔默研究实验室) ; Arizona State University(亚利桑那州立大学) ; Clemson University(克莱姆森大学)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
机构 * College of Computer Engineering, Jimei University(嘉应大学计算机工程学院) ; Chengyi College, Jimei University(嘉应大学 Chengyi 学院) ; Department of Technology, Management and Economics, Technical University of Denmark(丹麦技术大学技术、管理与经济系)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :harmlessness(abstract);分类 cs.CL、cs.LG
机构 * University of Udine, Italy(乌迪内大学)
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.LG
Comments Full version of the paper accepted for publication at the 28th European Conference on Artificial Intelligence (ECAI 2025)
机构 * Dongguan University of Technology(东莞科技学院) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments The paper has been accepted for presentation at ECAI 2025
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Department of Mechanical Engineering, Imperial College London, UK(帝国理工学院机械工程系) ; Department of Computing, Imperial College London, UK(帝国理工学院计算系) ; Department of Informatics, King's College London, UK(伦敦国王学院信息学系)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments 9 pages, 6 figures. Accepted for publication at The 14th Conference on Prestigious Applications of Intelligent Systems (PAIS-2025)
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY
Comments Accepted to AIES 2025
机构 * organization= UrbanResilience.AI Lab, Zachry Department of Civil ; Environmental Engineering, Texas A\&M University , city= College Station , postcode= 77843 , state= TX , country= USA
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.CY
机构 * Institute for Applied Informatics (InfAI) at Leipzig University(应用信息学院(InfAI))
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments to be published in the workshop proceedings of the "From Rules to Language Models: Comparative Performance Evaluation" workshop, held alongside RANLP 2025
机构 * School of Computer Science, Peking University(北京大学计算机科学学院) ; School of Information Science and Engineering, Chongqing Jiaotong University(重庆交通大学信息科学与工程学院) ; China Research and Development Academy of Machinery Equipment(机械电子研究发展院) ; Center of Information Research, Academy of Military Science(军事科学信息研究中心)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments 12 pages, 5figures
机构 * School of Electrical and Computer Engineering, University of Sydney(悉尼大学电气与计算机工程学院) ; School of Computer Science, University of Adelaide(阿德莱德大学计算机科学学院) ; School of Computing and Information Technology, University of Wollongong(沃林根大学计算与信息科技学院)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
机构 * University of California at Berkeley(加州大学伯克利分校) ; Indian Institute of Technology Bombay(印度班加罗尔理工学院) ; Chalmers University of Technology and University of Gothenburg(查尔姆斯理工大学和哥德堡大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments This work has been accepted at ATVA'25
机构 * Otto von Guericke University Magdeburg(奥托·冯·格里克大学马格德堡)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY
Comments Accepted to AAAI/ACM AIES 2025
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY
机构 * Scale AI
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG
机构 * NLP(自然语言处理) ; CT Lab, Department of Computer and Information Science, University of Macau(计算机与信息科学系,澳门大学) ; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院) ; Provable Responsible AI and Data Analytics Lab, KAUST(可证明责任AI与数据分析实验室,卡尔斯兰大学) ; Hong Kong Baptist University(香港 Baptist 大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
Comments Accepted to TACL 2025. This version is a pre-MIT Press publication version
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
机构 * Microsoft Industry AI(微软产业人工智能)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments Benchmark available at: https://huggingface.co/datasets/nogabenyoash/SecQue
Journal ref n Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics, Association for Computational Linguistics (2025) https://aclanthology.org/2025.gem-1.16/
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
Comments 8 pages, 5 figures
机构 * BUPT(北京邮电大学) ; WeChat Vision, Tencent Inc.(腾讯公司) ; Tsinghua University(清华大学)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
Comments Working in progress
专题命中 安全评测 :trustworthy(abstract);分类 cs.CL、cs.AI
机构 * Department of Computer Science and Engineering, BGC Trust University Bangladesh(Bangladesh BGC Trust 大学 计算机科学与工程系) ; Department of Computer Science and Engineering, International Islamic University Chittagong(Bangladesh 国际伊斯兰大学 昌德加荣分校 计算机科学与工程系) ; Centre for Securing Digital Futures, School of Science, Edith Cowan University(澳大利亚 埃德温·考文大学 科学学院 安全数字未来中心)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI、cs.LG