Policy Maps: Tools for Guiding the Unbounded Space of LLM Behaviors
机构 * Stanford University(斯坦福大学) ; Apple(苹果公司) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
Comments UIST 2025
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Stanford University(斯坦福大学) ; Apple(苹果公司) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 安全训练 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
Comments UIST 2025
机构 * Bosch Research North America(博世北美研究部) ; Bosch Center for Artificial Intelligence (BCAI)(博世人工智能中心) ; Texas A&M University(德克萨斯A&M大学)
专题命中 安全训练 :alignment(abstract);分类 cs.AI
专题命中 越狱攻击 :jailbreak(title,abstract)
机构 * Politecnico di Milano(米兰理工学院)
专题命中 提示注入 :prompt injection(title,abstract);jailbreak(abstract);分类 cs.AI、cs.LG
Comments 22 pages, preprint
机构 * CSE, HKUST(香港科技大学计算机科学与工程系) ; CSE, CUHK(香港城市大学计算机科学与工程系) ; Theory Lab, Huawei(华为理论实验室)
专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI
Comments 9 pages, preprint, code: https://github.com/HKUST-KnowComp/AutoSchemaKG
机构 * IBM Research(IBM研究院) ; Princeton University(普林斯顿大学)
专题命中 隐私与版权 :safety(abstract);分类 cs.CL、cs.AI
机构 * Amirkabir University of Technology(阿姆irkabir技术大学) ; Part AI Research Center(Part人工智能研究中心) ; University of Mazandaran(马赞德兰大学) ; King’s College London(伦敦国王学院)
专题命中 安全评测 :alignment(title);分类 cs.CL
Comments Preprint. Under review
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
机构 * University of Warwick(沃里克大学) ; Wuhan University of Technology(武汉理工大学)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
Comments 44 pages,11 figures
机构 * Department of Electrical and Computer Engineering, Western University(西方大学电气与计算机工程系) ; James Watt School of Engineering, University of Glasgow(格拉斯哥大学詹姆斯·瓦特工程学院) ; Department of Electrical and Computer Engineering, University of Waterloo(滑铁卢大学电气与计算机工程系)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Journal ref IEEE Network, 2025
机构 * Huazhong University of Science and Technology(华中科技大学) ; Lehigh University(莱斯大学) ; The University of Hong Kong(香港大学) ; Jilin University(吉林大学) ; Southern University of Science and Technology(南方科技大学) ; Worcester Polytechnic Institute(沃思堡理工学院) ; LinkedIn Corporation(领英公司) ; Squirrel Ai Learning ; University of Georgia(佐治亚大学) ; Duke University(杜克大学) ; Michigan State University(密歇根州立大学) ; Salesforce Research(Salesforce研究) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; University of Illinois at Chicago(伊利诺伊大学芝加哥分校) ; Microsoft Research(微软研究院)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.AI
Comments 87 pages, 21 figures, 9 tables
专题命中 AI治理与伦理 :safety(abstract);分类 cs.CY
Comments pre-print
专题命中 其他安全 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
机构 * Tencent Inc.(腾讯公司)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG
机构 * Cooperative Medianet Innovation Center, Shanghai Jiao Tong University(合作中位网创新中心,上海交通大学) ; School of Artificial Intelligence, Shanghai Jiao Tong University(人工智能学院,上海交通大学) ; Shanghai Innovation Institute(上海创新研究院) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Beijing Institute of Technology(北京理工大学) ; A*STAR Centre for Frontier AI Research(A*STAR前沿人工智能研究中心)
专题命中 其他安全 :alignment(abstract);分类 cs.AI
Comments 16 pages, Accepted at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
专题命中 其他安全 :alignment(abstract)
Comments 9 pages, 7 figures, 4 tables
专题命中 其他安全 :safety(abstract)
专题命中 其他安全 :alignment(abstract)
Comments 19 pages, 6 figures