Enhance the Safety in Reinforcement Learning by ADRC Lagrangian Methods
通过ADRC拉格朗日方法增强强化学习的安全性
Mingxu Zhang, Huicheng Zhang, Jiaming Ji, Yaodong Yang, Ying Sun
机构
*
AI Thrust, The Hong Kong University of Science and Technology (Guangzhou)(人工智能方向,香港科技大学(广州))
;
School of Artificial Intelligence, Peking University, Beijing, China(人工智能学院,北京大学,北京,中国)
;
Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家)
;
School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家)
机构
*
Zhejiang University(浙江大学)
;
National University of Singapore(新加坡国立大学)
;
Heriot-Watt University(赫瑞瓦特大学)
;
Southern University of Science and Technology(南方科技大学)
;
University of California, San Diego(加利福尼亚大学圣迭戈分校)
;
Northeastern University(东北大学)
CommentsInternational Conference on Machine Learning (ICML), 2026. v4: Revised CPC theory to (a) show enforced smoothness for likelihood-ratio control parameter and (b) for general control parameter, assume smoothness only of constrained policy $π_t^{(β)}$ w.r.t. $β$ (rather than of $(l_i - α)\cdotπ_t^{(β)}$ or of conformal weights)
`From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio LLMs
从提示到扰动:针对音频大语言模型的自适应语音越狱框架
Linghan Huang, Bo Li, Huaming Chen, Kim-Kwang Raymond Choo
机构
*
School of Electrical and Computer Engineering, The University of Sydney(悉尼大学电气与计算机工程学院)
;
University of Chicago(芝加哥大学)
;
University of Texas at San Antonio(德克萨斯大学圣安东尼奥分校)
Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment
回答前看清楚:通过显著性驱动的感知重新对齐减轻LVLMs中的幻觉
Pengxu Chen, Yao Zhu, Guangming Zhu, Jun Sheng, Jincai Huang, Xiangyang Ji, Liang Zhang
机构
*
Xidian University(西安电子科技大学)
;
Tsinghua University(清华大学)
;
Shanghai Road Transport Development Center(上海市道路运输发展中心)
;
Hunan Institute of Advanced Technology(湖南先进技术研究院)
Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation
M P V S Gopinadh
专题命中
安全评测
:safety(title,abstract);分类 cs.CL、cs.AI
Comments3 pages. Accepted at ACL 2026 Workshop on Evaluation in Practice: Methodological Rigor, Sociotechnical Perspectives, & Community Collaboration (EvalEval)
Orphan risks at the frontier of artificial intelligence: What diverging safety and compliance frameworks reveal about how AI companies choose the risks they prioritize