RADIANT: Retrieval AugmenteD entIty-context AligNmenT -- Introducing RAG-ability and Entity-Context Divergence
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 偏好对齐 :alignment(title,abstract);DPO(abstract);分类 cs.CL
机构 * Northeastern University, China(东北大学) ; Alibaba Group, Hangzhou, China(阿里巴巴集团)
专题命中 偏好对齐 :RLHF(abstract,comments);safety(abstract);分类 cs.CL、cs.CY
Comments Compared with the previous version, reinforcement learning has been added (as a new section), including RLHF, RLVR, and RLAIF
机构 * AiLECS Lab, Monash University Melbourne, Australia(墨尔本大学AiLECS实验室,澳大利亚) ; AiLECS Lab, Monash University, Australia ICMEC Australia, Sydney, Australia(墨尔本大学AiLECS实验室,澳大利亚 ICMEC澳大利亚,悉尼,澳大利亚)
专题命中 偏好对齐 :DPO(abstract);safety(abstract);分类 cs.AI、cs.LG
Comments Accepted for publication in the Proceedings of the 8th International Conference on Algorithms, Computing and Artificial Intelligence (ACAI 2025)
机构 * The Knowledge Engineering Group (KEG), Tsinghua University(清华大学知识工程小组(KEG)、清华大学)
专题命中 偏好对齐 :DPO(abstract);分类 cs.AI、cs.LG
Comments 21 pages, 4 figures
机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) ; School of Computer Science(计算机科学学院) ; Baidu Inc.(百度公司) ; School of Computer Science, Wuhan University(武汉大学计算机学院)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI
机构 * Zhejiang University(浙江大学) ; Soochow Securities Co.,Ltd.(苏州证券有限公司) ; HKUST(GZ)(香港科技大学(广州)) ; Nanjing University(南京大学)
专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI
Comments EMNLP 2025 Main
机构 * Yeshiva University(耶鲁大学) ; Oklahoma State University(俄克拉荷马州立大学) ; Case Western Reserve University(凯斯西储大学) ; North Carolina State University(北卡罗来纳州立大学)
专题命中 偏好对齐 :alignment(abstract);分类 cs.LG
机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) ; Institute of Artificial Intelligence, China Telecom (TeleAI)(中国电信人工智能研究院(TeleAI)) ; Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
专题命中 偏好对齐 :DPO(abstract);分类 cs.CL
机构 * UC San Diego(加州大学圣迭戈分校) ; Adobe Research(Adobe研究)
专题命中 偏好对齐 :DPO(abstract);分类 cs.LG
Comments 10 pages
机构 * SCIR Lab, Harbin Institute of Technology, China(哈尔滨工业大学SCIR实验室) ; City University of Hong Kong(香港城市大学) ; National University of Singapore(新加坡国立大学)
专题命中 安全训练 :safety(title,abstract);分类 cs.CL
机构 * Zhejiang University and Alibaba Cloud(浙江大学和阿里云) ; Alibaba Cloud(阿里云) ; Zhejiang University(浙江大学)
专题命中 安全训练 :alignment(title,abstract)
Comments Accepted by T-ASE and CoRL25 GenPriors Workshop
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Tsinghua University(清华大学) ; University of Science and Technology of China(中国科学技术大学)
专题命中 安全训练 :alignment(abstract);safety(abstract);分类 cs.LG
机构 * University of Electronic Science and Technology of China(电子科技大学) ; Huazhong University of Science and Technology(华中科技大学) ; The University of Queensland(昆士兰大学)
专题命中 越狱攻击 :alignment(title,abstract);safety(title,abstract);分类 cs.LG
Comments 16 pages,9 figures
机构 * Northwestern University(西北大学) ; University of Illinois at Chicago(伊利诺伊大学香槟分校)
专题命中 越狱攻击 :safety(title,abstract);jailbreak(abstract);分类 cs.AI;trustworthy(comments)
Comments accepted by the Trustworthy FMs workshop in ICCV 2025
机构 * Northeastern University(东北大学) ; Stevens Institute of Technology(斯蒂文斯理工学院)
专题命中 越狱攻击 :safety(abstract);jailbreak(abstract);分类 cs.AI
机构 * Ranchview High School(拉文斯维尔高中) ; Soongsil University(松山大学)
专题命中 提示注入 :prompt injection(title,abstract);jailbreak(abstract);分类 cs.AI
Comments 8 pages, 4 figures, 2 tables
专题命中 幻觉与事实性 :trustworthy(title);safety(abstract);分类 cs.AI、cs.CY、cs.LG
Comments PhD Thesis
机构 * Xin Tong School of Information and Cyber Security People’s Public Security University of China(信息与网络安全学院 中国人民公安大学) ; Zhi Lin School of Safety Science Tsinghua University(安全科学学院 清华大学) ; Jingya Wang School of Information and Cyber Security People’s Public Security University of China(信息与网络安全学院 中国人民公安大学) ; Bo Jin* The Third Research Institute of the Ministry of Public Security of China(公安部第三研究所)
专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI
机构 * Nankai University(南开大学) ; Tsinghua University(清华大学) ; Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) ; Peng Cheng Laboratory(鹏城实验室) ; Shenzhen ShenNong Information Technology Co., Ltd.(深圳深农信息技术有限公司)
专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI
Comments EMNLP 2025
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI
专题命中 安全评测 :safety(title,abstract);分类 cs.CY
Comments 7 pages, 3 figures, submitted to EMNLP 2025 and ECAT Research Workshop 2025
机构 * Department of Computer Engineering, Sharif University of Technology(谢里夫理工大学计算机工程系)
专题命中 安全评测 :alignment(abstract);safety(abstract);trustworthy(abstract);分类 cs.CL
专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.LG
机构 * Advanced Institute of So-Go-Chi (Convergence Knowledge) Informatics(融合知识研究院) ; Tohoku University(东北大学) ; Japan Advanced Institute of Science and Technology(日本先进科学研究院) ; Faculty of Information and Communication Technology(信息与通信技术学院) ; Mahidol University(玛希敦大学)
专题命中 安全评测 :trustworthy(title);分类 cs.AI
Journal ref Natural Language Processing and Information Systems (NLDB 2025)
专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.LG
Comments Accepted to ICLR 2025
机构 * Oxford Digital Health Labs(牛津数字健康实验室) ; Nuffield Department of Women’s and Reproductive Health(妇女与生殖健康尼富尔德部门) ; University of Oxford(牛津大学) ; OATML ; Department of Computer Science(计算机科学系)
专题命中 安全评测 :safety(abstract);分类 cs.CL、cs.AI
机构 * Universidad de León(莱昂大学)
专题命中 安全评测 :alignment(abstract);分类 cs.AI、cs.CY
Comments in Spanish language
机构 * Fujian University of Technology(福建工程学院) ; Fujian Provincial Key Laboratory of Big Data Mining and Applications(福建省大数据挖掘与应用重点实验室) ; Key Laboratory of Biomedical Imaging Science and System, Chinese Academy of Sciences(生物医学成像科学与系统重点实验室,中国科学院)
专题命中 安全评测 :alignment(abstract);分类 cs.AI
专题命中 安全评测 :alignment(abstract);分类 cs.LG
机构 * Radian Group Inc.(Radian集团) ; Sri Sivasubramaniya Nadar College Of Engineering(Sri Sivasubramaniya纳达尔工程学院) ; Amazon(亚马逊)
专题命中 安全评测 :trustworthy(abstract);分类 cs.AI
Comments 8 pages, 2 figures, Accepted at the OARS Workshop, KDD 2025, Paper link: https://oars-workshop.github.io/papers/Raman2025.pdf