Reward Model Routing in Alignment
机构 * National University of Singapore(新加坡国立大学)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * National University of Singapore(新加坡国立大学)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI
机构 * Bloomberg(高盛)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Oracle Health AI(Oracle健康AI)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Data and Web Science Group, University of Mannheim(曼海姆大学数据与网络科学组) ; Vision and AI Lab, Indian Institute of Science(印度科学院视觉与人工智能实验室) ; Carnegie Mellon University(卡内基梅隆大学) ; Max-Planck-Institute for Informatics, Saarland Informatics Campus(马克斯·普朗克信息研究所,萨尔兰信息校园)
专题命中 安全训练 :safety(title,abstract);分类 cs.CL、cs.LG
Comments Accepted at NeurIPS 2025 bWorkshop Lock-LLM. *Equal Contribution
专题命中 安全训练 :safety(title,abstract)
专题命中 安全训练 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
Comments preprint
机构 * Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学)
专题命中 安全训练 :safety(abstract);分类 cs.AI
专题命中 安全训练 :alignment(abstract);分类 cs.AI
机构 * Oregon State University(俄勒冈州立大学) ; University of British Columbia(不列颠哥伦比亚大学) ; University of Toronto(多伦多大学) ; George Mason University(乔治·梅森大学) ; Rutgers University(罗格斯大学) ; University of Texas at Arlington(德克萨斯大学阿灵顿分校)
专题命中 越狱攻击 :alignment(abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.LG
Comments Pre-print
机构 * ETH Zürich(苏黎世联邦理工学院) ; Stanford CS(斯坦福大学计算机科学系)
专题命中 越狱攻击 :alignment(abstract);jailbreak(abstract);分类 cs.AI、cs.LG
机构 * SRI ; The Ohio State University(俄亥俄州立大学)
专题命中 越狱攻击 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG
机构 * LIRIS - CNRS, INSA Lyon, Universite Claude Bernard Lyon 1(LIRIS - CNRS,INSA里昂,克劳德·贝尔纳大学里昂) ; Esker ; ServiceNow Research(ServiceNow研究) ; Mila - Quebec AI Institute(魁北克人工智能研究所) ; McGill University(麦吉尔大学) ; Polytechnique Montréal(蒙特利尔理工学院)
专题命中 越狱攻击 :prompt injection(abstract);分类 cs.CL
机构 * Language Technology Lab, University of Cambridge(剑桥大学语言技术实验室)
专题命中 幻觉与事实性 :alignment(abstract);trustworthy(abstract);分类 cs.CL
机构 * University of Utah(犹他大学) ; Monash University(莫纳什大学) ; Texas A&M University(德克萨斯农工大学)
专题命中 幻觉与事实性 :safety(abstract);trustworthy(abstract);分类 cs.AI
机构 * Princeton University(普林斯顿大学)
专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI
专题命中 隐私与版权 :trustworthy(abstract);分类 cs.AI
Comments 12 pages, 2 figures, 4 tables
机构 * Institute of Image Communication and Network Engineering(图像通信与网络工程研究所) ; MoE Key Lab of Artificial Intelligence, AI Institute(人工智能关键实验室) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 安全评测 :alignment(title,abstract)
机构 * Singapore Management University(新加坡管理大学)
专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.CL、cs.AI、cs.LG
Comments 16 pages, 4 figures
机构 * University of Chicago(芝加哥大学) ; University of Illinois(伊利诺伊大学) ; Virtue AI ; Meta
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG
Comments 60 pages, 16 figures
机构 * University of Virginia(弗吉尼亚大学) ; Dexcom(德科姆公司)
专题命中 安全评测 :safety(abstract);分类 cs.AI、cs.CY、cs.LG
机构 * Iowa State University(爱荷华州立大学)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.LG
机构 * University of Mannheim(曼海姆大学) ; GESIS - Leibniz Institute for the Social Sciences(莱布尼茨社会科学研究所) ; Complexity Science Hub Vienna(维也纳复杂科学中心)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
Comments Accepted to EMNLP Findings 2025
机构 * Institute for Machine Vision, University of Applied Sciences Kempten(机器视觉研究所,应用科技大学凯普腾)
专题命中 安全评测 :alignment(abstract);分类 cs.LG
Comments This work has been submitted to the IEEE for possible publication
专题命中 安全评测 :safety(abstract);分类 cs.LG
Comments 40 pages, 19 figures, 9 tables
机构 * East China University of Science and Technology(东华大学) ; Singapore Management University(新加坡管理学院)
专题命中 安全评测 :alignment(abstract)
机构 * Laboratory of Multimodal Research In Industry, AI Institute, Innopolis University(工业多模态研究实验室,人工智能研究所,因诺普利斯大学) ; Phystech School of Applied Mathematics and Computer Science, Moscow Institute of Physics and Technology(物理与技术莫斯科应用数学与计算机科学学院,莫斯科物理技术学院) ; Research Center for Artificial Intelligence, Innopolis University(人工智能研究中心,因诺普利斯大学) ; Q Deep, Innopolis(Q深度,因诺普利斯) ; Machine Learning and Data Representation Lab, Innopolis University(机器学习与数据表示实验室,因诺普利斯大学)
专题命中 安全评测 :safety(abstract)
Comments 12 pages, 5 figures, 2 tables, ICOMP 2025
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
专题命中 安全评测 :alignment(abstract)
机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) ; ARC Lab, Tencent PCG(腾讯PCG ARC实验室) ; MAIS, Institute of Automation, CAS, Beijing(自动化研究所北京研究所MAIS)
专题命中 安全评测 :alignment(abstract)
专题命中 安全评测 :safety(abstract)
Comments Accepted to IEEE TPAMI, extension of NeurIPS'23 work: A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)
机构 * Vrije University Amsterdam(荷兰阿姆斯特丹自由大学) ; Tri-institutional Center for Translational Research in Neuroimaging(转化神经影像研究联合中心) ; Emory University(埃默里大学) ; Key Laboratory of Genetic Evolution and Animal Models(遗传进化与动物模型重点实验室) ; Kunming Institute of Zoology(昆明动物研究所) ; Chinese Academy of Sciences Kunming(中国科学院昆明分院) ; Department of Psychiatry, Amsterdam UMC, University of Amsterdam(阿姆斯特丹大学精神病科) ; Department of Physics and Technology, UiT The Arctic University of Norway(北极大学挪威理工学院物理与技术系)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG
Comments This manuscript has been accepted by Biomedical Signal Processing and Control and the code is available at https://github.com/TianzhengHU/BrainIB_coding/tree/main/BrainIB_GIB