On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization
Jiancong Xiao, Ziniu Li, Xingyu Xie, Emily Getzen, Cong Fang, Qi Long, Weijie J. Su
机构
*
University of Pennsylvania(宾夕法尼亚大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
National University of Singapore(新加坡国立大学)
;
Peking University(北京大学)
;
Joint corresponding authors(联合通讯作者)
KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
Jason R Brown, Lennie Wells, Edward James Young, Sergio Bacallado
机构
*
Computational and Biological Learning Group, Department of Engineering, University of Cambridge, Cambridge, UK(计算生物学学习组,工程系,剑桥大学,剑桥,英国)
;
Department of Computer Science and Technology, University of Cambridge, Cambridge, UK(计算机科学与技术系,剑桥大学,剑桥,英国)
;
Statistics Laboratory, Department of Pure Mathematics and Mathematical Statistics, University of Cambridge, UK(统计实验室,纯粹数学与数学统计系,剑桥大学,英国)
机构
*
New Laboratory of Pattern Recognition, CASIA(模式识别新实验室,中国科学院自动化研究所)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Nanjing university(南京大学)
;
Academy of Broadcasting Science, NRTA(广播科学研究院,国家广播电视总局)
WST: Weak-to-Strong Knowledge Transfer via Reinforcement Learning
Haosen Ge, Shuo Li, Lianghuan Huang
机构
*
Wharton AI & Analytics Initiative(沃顿人工智能与分析倡议)
;
University of Pennsylvania(宾夕法尼亚大学)
;
Department of Computer and Information Science(计算机与信息科学系)
;
Department of Physics and Astronomy(物理学与天文学系)
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
Baidu Inc.(百度公司)
EMPOWER: Evolutionary Medical Prompt Optimization With Reinforcement Learning
Yinda Chen, Yangfan He, Jing Yang, Dapeng Zhang, Zhenlong Yuan, Muhammad Attique Khan, Jamel Baili, Por Lip Yee
机构
*
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知国家重点实验室,中国科学技术大学)
;
Department of Computer Science, University of Minnesota-Twin Cities(计算机科学系,明尼苏达大学双城分校)
;
Center of Research for Cyber Security and Network (CSNET), Faculty of Computer Science and Information Technology, Universiti Malaya(网络安全与网络研究中心(CSNET),马来亚大学计算机科学与信息技术学院)
;
DSLAB, School of Information Science & Engineering, Lanzhou University(信息科学与工程学院,兰州大学)
;
Institute of Computing Technology, Chinese Academy of Sciences(计算技术研究所,中国科学院)
;
Department of AI, Prince Mohammad bin Fahd University(人工智能系,普林姆·法赫德大学)
;
Department of Computer Engineering, College of Computer Science, King Khalid University(计算机工程系,国王·卡利德大学)
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation
Jun Zhuang, Haibo Jin, Ye Zhang, Zhengjian Kang, Wenbin Zhang, Gaby G. Dagher, Haohan Wang
机构
*
Boise State University(博伊州立大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Pittsburgh(匹兹堡大学)
;
New York University(纽约大学)
;
Florida International University(佛罗里达国际大学)
CommentsAccepted for EMNLP'25 Findings. TL;DR: We propose a new two-stage intent-based prompt-refinement framework, IntentPrompt, that aims to explore the vulnerability of LLMs' content moderation guardrails by refining prompts into benign-looking declarative forms via intent manipulation for red-teaming purposes
An Outlook on the Opportunities and Challenges of Multi-Agent AI Systems
Fangqiao Tian, An Luo, Jin Du, Xun Xian, Robert Specht, Ganghua Wang, Xuan Bi, Jiawei Zhou, Ashish Kundu, Jayanth Srinivasa, Charles Fleming, Rui Zhang, Zirui Liu, Mingyi Hong, Jie Ding
CommentsPublished at Transaction of Machine Learning Research 08/2025, Large Language Models (LLMs), Interference-time activation shifting, Steerability, Explainability, AI alignment, Interpretability
L-XAIDS: A LIME-based eXplainable AI framework for Intrusion Detection Systems
Aoun E Muhammad, Kin-Choong Yow, Nebojsa Bacanin-Dzakula, Muhammad Attique Khan
专题命中
越狱攻击
:safety(abstract);分类 cs.AI
CommentsThis is the authors accepted manuscript of an article accepted for publication in Cluster Computing. The final published version is available at: 10.1007/s10586-025-05326-9
机构
*
Department of Electrical and Computer Engineering, University of Victoria(电气与计算机工程系,维多利亚大学)
;
Department of Mechanical Engineering, University of Victoria(机械工程系,维多利亚大学)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
Adi Simhi, Itay Itzhak, Fazl Barez, Gabriel Stanovsky, Yonatan Belinkov
机构
*
Technion – Israel Institute of Technology(技术ion-以色列理工学院)
;
University of Oxford and WhiteBox(牛津大学和WhiteBox)
;
School of Computer Science and Engineering, The Hebrew University of Jerusalem(耶路撒冷希伯来大学计算机科学与工程学院)