Sparse Rewards Can Self-Train Dialogue Agents
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL
Comments Accepted to ACL 2025 (Findings)
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL
Comments Accepted to ACL 2025 (Findings)
机构 * Meta FAIR ; University of Washington(华盛顿大学) ; Meta GenAI
专题命中 偏好对齐 :trustworthy(abstract);分类 cs.AI
Comments 17 pages, 6 tables, 5 figures
机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; Renmin University of China(中国人民大学)
专题命中 越狱攻击 :safety(title,abstract);alignment(abstract);分类 cs.CL
机构 * Department of Electrical Engineering(电气工程系) ; Indian Institute of Technology Delhi(印度理工学院德里)
专题命中 越狱攻击 :jailbreak(abstract);分类 cs.CL
机构 * TIB – Leibniz Information Centre for Science and Technology(莱比锡信息科学与技术研究中心) ; Amazon, Germany(亚马逊公司)
专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Google DeepMind(谷歌DeepMind) ; Google Research(谷歌研究)
专题命中 安全评测 :alignment(title,abstract);safety(title,abstract);分类 cs.AI、cs.LG
Comments 28 pages, 16 figures
机构 * Professorship of Autonomous Vehicle Systems, TUM School of Engineering and Design, Technical University of Munich(自主车辆系统教授职位,技术大学慕尼黑工程与设计学院) ; Munich Institute of Robotics and Machine Intelligence (MIRMI)(慕尼黑机器人与机器智能研究所) ; TUM School of Computation, Information and Technology, Department of Computer Engineering, Technical University of Munich(技术大学慕尼黑计算、信息与技术学院,计算机工程系)
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI
Comments Final Version and Paper Accepted at IEEE ITSC 2025
机构 * organization= Orion Lab, National Observatory of Athens \& National Technical University of Athens , country= Greece ; organization= Department of Informatics \& Telematics, Harokopio University of Athens , country= Greece ; organization= Faculty of Electrical Engineering ; organization= BIFOLD - Berlin Institute for the Foundations of Learning
专题命中 安全评测 :alignment(title,abstract)
Comments Accepted at the ISPRS Journal of Photogrammetry and Remote Sensing. Our code implementation and weights for all experiments are publicly available at https://github.com/Orion-AI-Lab/MindTheModalityGap
机构 * Toulouse School of Economics, Centre National de la Recherche Scientifique (TSM-R), Université Toulouse Capitole(图卢兹经济学院,法国国家科学研究中心(TSM-R),图卢兹大学)
专题命中 安全评测 :trustworthy(title);分类 cs.AI、cs.CY
机构 * The Polytechnic School, Ira A. Fulton Schools of Engineering, Arizona State University(亚利桑那州立大学工程学院Polytechnic学院) ; School of Manufacturing Systems and Networks, Ira A. Fulton Schools of Engineering, Arizona State University(亚利桑那州立大学工程学院制造系统与网络学院)
专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI
机构 * Ufonia Limited(乌菲尼亚有限公司) ; University of York(约克大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments 29 pages
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.LG
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments NeurIPS 2024 Workshop on Adaptive Foundation Models
机构 * School of Computing(计算学院) ; George Mason University(乔治·玛莎大学) ; College of Science(科学学院)
专题命中 安全评测 :safety(abstract);分类 cs.AI
Journal ref IEEE BigData, Year: 2024; Page: 3258-3263
机构 * Department of Artificial Intelligence, Korea University(人工智能系,韩国大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL
专题命中 安全评测 :safety(abstract);分类 cs.AI
Comments 38 pages;12 figures;12 tables
机构 * Meta
专题命中 安全评测 :safety(abstract);分类 cs.AI
机构 * Eindhoven University of Technology(埃因霍温理工大学)
专题命中 安全评测 :alignment(abstract)
Comments accepted at ICCV'25 workshop CV4BIOM
专题命中 安全评测 :trustworthy(abstract)
Comments TOSEM 2030 Special Issue
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CY
Comments 44 pages
机构 * ITU Artificial Intelligence and Data Science Research Center(伊斯坦布尔技术大学人工智能与数据科学研究中心) ; Department of Aeronautical Engineering(航空工程系) ; Eatron Technologies(Eatron技术公司) ; Department of Electronics and Communication Engineering(电子与通信工程系) ; ITU Artificial Intelligence and Data Science Application and Research Center(伊斯坦布尔技术大学人工智能与数据科学应用与研究中心) ; Department of Computer Engineering(计算机工程系)
专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG
Comments 7 pages, 7 figures, 2 tables, published in IEEE International Conference on Robotics and Automation (ICRA), June 2, 2023, London, UK
Journal ref IEEE International Conference on Robotics and Automation (ICRA), 2023, pp. 5631-5637
专题命中 其他安全 :alignment(title,abstract)
Comments This work has been received by CDC2025
机构 * Department of Computer Science and Information Systems, Birla Institute of Technology and Science, Pilani, Hyderabad Campus(计算机科学与信息系统系,比拉理工学院,比拉理工学院,海得拉巴校区)
专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG
机构 * Microsoft Research, Health Futures, Cambridge, United Kingdom(微软研究院,健康未来,剑桥,英国)
专题命中 其他安全 :safety(abstract);分类 cs.LG
Comments Actionable Interpretability Workshop at ICML 2025. 24 pages, 7 figures, 5 tables
专题命中 其他安全 :safety(abstract)