PIPA: Preference Alignment as Prior-Informed Statistical Estimation
机构 * The University of Texas at Austin, US(德克萨斯大学奥斯汀分校)
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.AI、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * The University of Texas at Austin, US(德克萨斯大学奥斯汀分校)
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.AI、cs.LG
机构 * Aalto University(阿alto大学) ; ETH Zürich(苏黎世联邦理工学院) ; KTH Royal Institute of Technology(皇家理工学院)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(abstract);分类 cs.AI
专题命中 越狱攻击 :safety(title,abstract);jailbreak(abstract);分类 cs.CL
机构 * eBRAIN Lab, New York University Abu Dhabi (NYUAD), UAE(eBRAIN实验室,纽约大学阿布扎赫尔分校(NYUAD),阿联酋) ; AI and Digital Science Research Center, Technology Innovation Institute (TII), Abu Dhabi, UAE(人工智能与数字科学研究中心,技术创新研究所(TII),阿布扎赫尔,阿联酋)
专题命中 越狱攻击 :safety(abstract);分类 cs.LG
Comments 8 pages, 8 figures, 1 table
机构 * University of Queensland(昆士兰大学)
专题命中 隐私与版权 :alignment(abstract);分类 cs.AI、cs.LG
Comments The camera ready version for ECAI-2025
机构 * Kempelen Institute of Intelligent Technologies(凯普勒智能技术研究所) ; University of Copenhagen(哥本哈根大学) ; Comenius University in Bratislava(布拉迪斯拉瓦科เมนius大学)
专题命中 安全评测 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI、cs.CY
Comments ACL 2025 main
Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Volume 1: Long Papers)
机构 * Fudan University(复旦大学) ; Shanghai AI Lab(上海人工智能实验室) ; Tsinghua University(清华大学) ; The University of Hong Kong(香港大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
Comments Work in progress
机构 * National University of Singapore(新加坡国立大学)
专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Helmholtz AI(海德堡人工智能研究所) ; Helmholtz Munich(海德堡慕尼黑)
专题命中 安全评测 :trustworthy(abstract);分类 cs.LG
机构 * Department of Electronics, Information and Bioengineering, Politecnico di Milano(电子、信息与生物工程学院,米兰理工学院) ; Fondazione IRCCS Istituto Nazionale dei Tumori di Milano(米兰国家肿瘤研究所) ; Department of Medicine, Section of Hematology/Oncology, University of Chicago(医学学院,血液学/肿瘤学部门,芝加哥大学) ; LEARNLab, IRCCS Istituto Neurologico Carlo Besta(LEARN实验室,卡尔·贝斯塔神经病学研究所)
专题命中 安全评测 :trustworthy(abstract);分类 cs.LG
Comments Emilia Ambrosini and Simona Ferrante equally contributed to the work
机构 * Harvard College(哈佛学院)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
机构 * University of Toronto Department of Psychiatry(多伦多大学精神病学系) ; Toronto Metropolitan University(多伦多 Metropolitan 大学) ; St. Michael’s Hospital, Unity Health Toronto(圣米歇尔医院,统一健康多伦多) ; Department of Electrical, Computer, and Biomedical Engineering(电气、计算机和生物医学工程系) ; University of Waterloo(滑铁卢大学) ; Department of Management Science and Engineering(管理科学与工程系)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
专题命中 安全评测 :safety(abstract)
Comments arXiv admin note: text overlap with arXiv:2412.13528
专题命中 安全评测 :alignment(abstract)
Comments Accepted at ICCV 2025 (main conference)
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Rice University(Rice大学) ; Carnegie Mellon University(卡内基梅隆大学) ; AIFARMS ; Center for Digital Agriculture at UIUC(伊利诺伊大学厄巴纳-香槟分校数字农业中心)
专题命中 安全评测 :trustworthy(abstract)
Comments Project Website: https://agmmu.github.io/ Huggingface: https://huggingface.co/datasets/AgMMU/AgMMU_v1/
专题命中 其他安全 :alignment(title,abstract);safety(title,abstract)
机构 * Department of Civil Engineering(土木工程系) ; Department of Aerospace Engineering(航空航天工程系) ; Department of Civil, Environmental, and Geo-Engineering(土木、环境与地球工程系)
专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG
Journal ref "RACER: Rational Artificial Intelligence Car-Following-Model Enhanced by Reality," in IEEE Transactions on Intelligent Transportation Systems,
专题命中 其他安全 :alignment(abstract);分类 cs.CL
Journal ref Nature Machine Intelligence volume 7, pages1154-1167 (2025)
机构 * Hanyang University(翰阳大学)
专题命中 其他安全 :alignment(abstract)