机构
*
School of Computing and Information Systems, The University of Melbourne(计算与信息系统学院,墨尔本大学)
;
Melbourne Data Analytics Platform, The University of Melbourne(墨尔本数据分析平台,墨尔本大学)
;
Faculty of Information Technology, Monash University(信息技术学院,墨尔本大学)
Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
利用重要性采样将对齐模块从大语言模型中分离出来
Yi Liu, Dianqing Liu, Mingye Zhu, Junbo Guo, Yongdong Zhang, Zhendong Mao
机构
*
State Key Laboratory of Communication Content Cognition, People’s Daily Online(通信内容认知国家重点实验室,人民在线)
;
University of Science and Technology of China(中国科学技术大学)
Preference Orchestrator: Prompt-Aware Multi-Objective Alignment for Large Language Models
Biao Liu, Ning Xu, Junming Yang, Xin Geng
机构
*
School of Computer Science and Engineering(计算机科学与工程学院)
;
Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications(新一代人工智能技术及其交叉应用关键实验室)
;
Ministry of Education(教育部)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
Amr Gomaa, Ahmed Salem, Sahar Abdelnabi
机构
*
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI))
;
Microsoft(微软)
;
ELLIS Institute Tübingen and MPI for Intelligent Systems(图宾根ELLIS研究所和智能系统研究所)
;
Tübingen AI Center(图宾根人工智能中心)
Erhan Xu, Kai Ye, Hongyi Zhou, Luhan Zhu, Francesco Quinzan, Chengchun Shi
机构
*
Department of Statistics(统计系)
;
LSE London, UK(伦敦大学学院)
;
Department of Mathematics(数学系)
;
Tsinghua University(清华大学)
;
School of Design LCC, UAL London, UK(伦敦艺术大学设计学院)
;
Department of Engineering Science(工程科学系)
;
University of Oxford(牛津大学)
机构
*
Electrical and Computer Engineering University of Virginia(电气与计算机工程大学弗吉尼亚大学)
;
Princeton Language and Intelligence Princeton University(普林斯顿语言与智能普林斯顿大学)
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
Hongbin Zhang, Kehai Chen, Xuefeng Bai, Yang Xiang, Min Zhang
机构
*
Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen, China(计算与智能研究所,哈尔滨工业大学,深圳,中国)
;
Peng Cheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国)
A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment
Jean-Philippe Corbeil, Amin Dada, Jean-Michel Attendu, Asma Ben Abacha, Alessandro Sordoni, Lucas Caccia, François Beaulieu, Thomas Lin, Jens Kleesiek, Paul Vozila
KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
Jason R Brown, Lennie Wells, Edward James Young, Sergio Bacallado
机构
*
Computational and Biological Learning Group, Department of Engineering, University of Cambridge, Cambridge, UK(计算生物学学习组,工程系,剑桥大学,剑桥,英国)
;
Department of Computer Science and Technology, University of Cambridge, Cambridge, UK(计算机科学与技术系,剑桥大学,剑桥,英国)
;
Statistics Laboratory, Department of Pure Mathematics and Mathematical Statistics, University of Cambridge, UK(统计实验室,纯粹数学与数学统计系,剑桥大学,英国)
Cultivating Helpful, Personalized, and Creative AI Tutors: A Framework for Pedagogical Alignment using Reinforcement Learning
Siyu Song, Wentao Liu, Ye Lu, Ruohua Zhang, Tao Liu, Jinze Lv, Xinyun Wang, Aimin Zhou, Fei Tan, Bo Jiang, Hao Hao
机构
*
Shanghai Innavation Institute(上海创新研究院)
;
Shanghai Institute of AI for Education(上海人工智能教育研究院)
;
School of Computer Science and Technology(计算机科学与技术学院)
;
Department of Educational Information Technology(教育信息技术系)