LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling
机构 * Soochow University(苏州大学) ; LCM Laboratory(LCM实验室)
专题命中 偏好对齐 :safety(abstract);分类 cs.CL、cs.AI
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * Soochow University(苏州大学) ; LCM Laboratory(LCM实验室)
专题命中 偏好对齐 :safety(abstract);分类 cs.CL、cs.AI
机构 * Dept. of Computer Engineering(计算机工程系) ; SIES Graduate School of Technology(SIES技术研究生学院) ; Bharatiya Vidya Bhavan's Sardar Patel Institute of Technology(巴哈里亚·维达·巴万学院萨达尔·帕特尔技术学院)
专题命中 安全训练 :alignment(title,abstract);分类 cs.LG
机构 * Department of Engineering Science, University of Oxford(工程科学系,牛津大学)
专题命中 安全训练 :safety(abstract);分类 cs.LG
Comments Accepted at NeurIPS 2025
机构 * Penn State University(宾夕法尼亚州立大学) ; Information Sciences Institute, USC(信息科学研究所)
专题命中 越狱攻击 :jailbreak(title,abstract);alignment(abstract);safety(abstract);分类 cs.CL
专题命中 越狱攻击 :safety(abstract);分类 cs.AI
机构 * Clemson University(克莱姆森大学)
专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI、cs.CY
机构 * Stanford University(斯坦福大学) ; Columbia University(哥伦比亚大学) ; Kaggle ; University of Washington(华盛顿大学) ; NYU(纽约大学)
专题命中 隐私与版权 :alignment(title);分类 cs.LG
专题命中 隐私与版权 :alignment(abstract)
机构 * Department of Electrical Engineering and Computer Science, University of Stavanger, Norway(斯瓦尔巴大学电气工程与计算机科学系) ; School of Physics, Engineering and Technology, University of York, UK(约克大学物理、工程与技术学院)
专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI
Comments 22 Pages
专题命中 安全评测 :alignment(title);分类 cs.CL
机构 * Carnegie Mellon University(卡内基梅隆大学)
专题命中 安全评测 :alignment(title);分类 cs.LG
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI
机构 * The University of Tokyo(东京大学)
专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.CL
Comments ICCV2025 Workshop
机构 * American University of Beirut(美国贝鲁特美国大学) ; SogetiLabs Research and Innovation(SogetiLabs研究与创新) ; National Center for Remote Sensing(远程 sensing 国家中心)
专题命中 安全评测 :safety(abstract);trustworthy(abstract)
机构 * Department of Computer Science(计算机科学系) ; Tulane University(Tulane 大学) ; Department of Industrial Engineering and Management(工业工程与管理系) ; Aalto University(Aalto 大学)
专题命中 安全评测 :safety(abstract);分类 cs.CL
Comments Advances in Neural Information Processing Systems 2025 (NeurIPS 2025), Poster, https://neurips.cc/virtual/2025/loc/san-diego/poster/121400
专题命中 安全评测 :safety(abstract)
机构 * NYU Shanghai(纽约大学上海校区) ; New York University(纽约大学) ; University of Washington(华盛顿大学) ; MBZUAI ; Microsoft(微软) ; UIUC(伊利诺伊大学香槟分校)
专题命中 AI治理与伦理 :alignment(title,abstract);分类 cs.CL、cs.AI
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);jailbreak(abstract);AI safety(abstract)
机构 * RockCyber ; Kleiner Perkins ; Qorvex Consulting & Roshan Consulting ; SAS Institute ; OWASP ; Wentworth Institute of Higher Education & Machine Learning Professional ; Skylink Antenna ; Roshan Consulting & Robotic Process Automation ; Stanford University
专题命中 AI治理与伦理 :red teaming(abstract);分类 cs.AI
机构 * Centre for Language Studies, Centre for Language and Speech Technology, Radboud University,Nijmegen, The Netherlands(语言研究所以及语言与语音技术中心,拉德堡德大学,尼姆egen,荷兰) ; Behavioural Science Institute, Radboud University, Nijmegen, The Netherlands(行为科学研究所,拉德堡德大学,尼姆egen,荷兰)
专题命中 其他安全 :alignment(title,abstract);分类 cs.CL
Comments This preprint has not undergone peer review or any post-submission improvements or corrections. The Version of Record of this contribution is published in TSD 2025. Lecture Notes in Computer Science, vol 16029
机构 * University of California, Los Angeles(加州大学洛杉矶分校)
专题命中 其他安全 :safety(title,abstract);分类 cs.AI
专题命中 其他安全 :alignment(title)
Comments In the 16th International Conference on Internetware 2025. 13 pages
Journal ref Proceedings of the 16th International Conference on Internetware (2025) 413-425
专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG
机构 * Soochow University(苏州大学) ; LCM Laboratory(长文实验室) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Apple(苹果公司)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI
机构 * MPI-SWS(马克斯·普朗克所际研究所)
专题命中 其他安全 :safety(abstract);分类 cs.LG
Comments NeurIPS'25 paper
专题命中 其他安全 :alignment(abstract);分类 cs.AI
机构 * University of Texas at Arlington(德克萨斯理工大学) ; Johnson & Johnson Innovative Medicine(强生创新医学)
专题命中 其他安全 :alignment(abstract);分类 cs.LG
Journal ref NeurIPS 2025 2nd Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences
机构 * Ubiquitous Knowledge Processing Lab (UKP Lab)(通用知识处理实验室) ; Department of Computer Science(计算机科学系) ; Hessian Center for AI (hessian.AI)(黑森人工智能中心)
专题命中 其他安全 :alignment(abstract);分类 cs.CL
Comments EMNLP 2025 Findings
专题命中 其他安全 :alignment(abstract)