MuCPT: Music-related Natural Language Model Continued Pretraining
机构 * Tsinghua University(清华大学) ; Tencent Inc(腾讯公司)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Tsinghua University(清华大学) ; Tencent Inc(腾讯公司)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
机构 * College of Engineering University of Georgia(工程学院 乔治亚大学)
专题命中 安全评测 :safety(abstract);分类 cs.AI
Comments 23 pages, 4 figures, 3 tables
机构 * Ph.D. Student, Dept. of Civil Eng., University of Texas at Arlington. E-mail ; Assistant Professor, Dept. of Computer Sci. \& Eng., University of Texas at Arlington. E-mail ; Assistant Professor, Dept. of Civil Eng., University of Texas at Arlington. E-mail
专题命中 安全评测 :safety(abstract);分类 cs.AI
机构 * Centre for Smart Health, School of Nursing, The Hong Kong Polytechnic University(智能健康研究中心、护理学院、香港理工大学) ; Department of Language Science and Technology, The Hong Kong Polytechnic University(语言科学与技术系、香港理工大学) ; Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(电子与电气工程系、香港理工大学)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Comments 9 pages, 2 figures
专题命中 安全评测 :alignment(abstract);分类 cs.CL
机构 * Institute for Intelligent Systems(智能系统研究所) ; Esslingen University of Applied Sciences(应用科学大学埃斯林根) ; Department of Computer Science(计算机科学系) ; University of Freiburg(弗赖堡大学)
专题命中 安全评测 :safety(abstract)
机构 * Peking University(北京大学) ; ByteDance(字节跳动) ; Princeton University(普林斯顿大学) ; CASIA(中国科学院自动化研究所) ; The University of Chicago(芝加哥大学)
专题命中 安全评测 :alignment(abstract)
Comments Project Page: https://tyfeld.github.io/mmadaparellel.github.io/
机构 * University of Southern California(南加州大学) ; Amazon AGI(亚马逊人工智能实验室)
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG
机构 * Barclays, Model Risk Management(巴克莱银行,模型风险管理部门) ; Columbia University(哥伦比亚大学)
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL
Comments AAAI 2026 Oral. 14 pages (including appendix), 11 figures. Code, data, results, and additional resources are available at: https://model-editing.github.io
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI
Comments 12 pages, 4 figures, 1 table, includes Supplementary Materials, simulation code on GitHub (https://github.com/AerisSpace/SecondLawIntelligence )
机构 * Department of Computer Science and Engineering, BRAC University(计算机科学与工程系,BRAC大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG
机构 * Cornell Tech(康奈尔科技) ; Microsoft Research(微软研究院)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL
Comments To be published in the Proceedings of the 40th Annual AAAI Conference on Artificial Intelligence (AAAI 2026 Special Track on AI for Social Impact )
机构 * Microsoft Corporation(微软公司)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI
机构 * uni-potsdam(波恩大学)
专题命中 其他安全 :safety(title,abstract)
Comments Preprint
机构 * IT & Engineering International University of Applied Sciences(IT与工程国际应用科学大学)
专题命中 其他安全 :alignment(title,abstract)
Comments 11 pages, 3 figures
机构 * Team exalsius(exalsius团队) ; Technische Universität Berlin(柏林技术大学)
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
Comments 8pages, 3 figures, technical report
机构 * AI VIETNAM(AI越南) ; Carnegie Mellon University(卡内基梅隆大学) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; PTIT ; Northwestern University(西北大学)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
Comments Acccepted in MICCAI Workshop 2025
机构 * Leibniz University Hannover(莱布尼茨汉诺威大学)
专题命中 其他安全 :safety(abstract);分类 cs.LG
Comments Accepted for publication in IEEE Transactions on Robotics (T-RO) 2025
机构 * Sydney AI Center, The University of Sydney(悉尼人工智能中心,悉尼大学) ; CSIRO, Data61(澳大利亚联邦科学与工业研究组织、Data61) ; Shanghai Jiao Tong University(上海交通大学) ; City University of Macau(澳门城市大学) ; CASIA(中国科学院自动化研究所)
专题命中 其他安全 :safety(abstract);分类 cs.LG
Comments 12 pages, 5 figures
专题命中 其他安全 :alignment(abstract);分类 cs.AI
Comments Accepted at AAAI 2026 oral
专题命中 其他安全 :safety(abstract)
机构 * Laboratoire d'Informatique Paris Descartes (LIPADE), Université Paris Cité (France)(巴黎笛卡尔大学信息学实验室(LIPADE),巴黎城市大学(法国))
专题命中 其他安全 :alignment(abstract)
Comments This work has been submitted to the IEEE ISBI for possible publication