Exploring Multilingual Concepts of Human Value in Large Language Models: Is Value Alignment Consistent, Transferable and Controllable across Languages?
Comments11 pages, 7 figures. For code see https://github.com/phelps-sg/llm-cooperation Updated with minor corrections: - corrected typo: "mesa-optimiser" instead of "meso-optimiser" - Cited Yang et al (2023) in support of claim that LLMs can solve optimisation problems - Acknowledged Seth Aslin for corrections
机构
*
Guangdong Provincial Key Laboratory of Brain-inspired Intelligent Computation, Department of Computer Science and Engineering, Southern University of Science and Technology(广东省脑启发智能计算重点实验室,计算机科学与工程系,南方科技大学)
机构
*
Nanyang Technological University(南洋理工大学)
;
National University of Singapore(新加坡国立大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
A*STAR(科技研究局)
;
Southern University of Science and Technology(南方科技大学)
;
University of Science and Technology of China(中国科学技术大学)
;
The Pennsylvania State University(宾夕法尼亚州立大学)
;
TeleAI
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Zhejiang University(浙江大学)
;
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
Renmin University of China(中国人民大学)
;
University of California, San Diego(加州大学圣地亚哥分校)
;
Tencent(腾讯)
;
Georgia Institute of Technology(佐治亚理工学院)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
CommentsShort version accepted as a Tiny Paper at the International Conference on Learning Representations (ICLR) 2024. Long version accepted to the Conference on Empirical Methods in Natural Language Processing (EMNLP) 2024 Findings
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
我难以相信它不稳健:在嵌入漂移下安全分类器的灾难性崩溃
Subramanyam Sahoo, Vinija Jain, Divya Chaudhary, Aman Chadha
机构
*
Independent(独立研究者)
;
Meta AI
;
AWS Generative AI Innovation Center, Amazon Web Services(AWS生成式AI创新中心,亚马逊网络服务)
;
Northeastern University, Seattle, WA, USA(东北大学,西雅图,华盛顿州,美国)
;
Stanford University(斯坦福大学)
Progressive Multimodal Alignment for Continual Instruction Tuning
用于持续指令微调的渐进式多模态对齐
Duzhen Zhang, Yahan Yu, Qiaoyi Su, Jiahua Dong, Tielin Zhang
机构
*
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences(中国科学院脑科学与智能技术卓越创新中心)
;
Kyoto University(京都大学)
;
Migu Culture Technology Co.,Ltd.(咪咕文化科技有限公司)
;
State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与类脑智能技术国家重点实验室)
I'm Sorry Driver, I'm Afraid I Can't Do That: Appraising the Safety of LLMs within Automotive Contexts
抱歉,司机,恐怕我不能这么做:评估LLMs在汽车环境中的安全性
Shaun Feakins, Ibrahim Habli, Kim Littler, Robert Palin
机构
*
UKRI AI Centre for Doctoral Training in Safe Artificial Intelligence Systems (SAINTS)(英国研究理事会安全人工智能系统博士培训中心(SAINTS))
;
University of York(约克大学)
;
Jaguar Land Rover(捷克·陆罗恩)