ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
ParaSpeechCLAP:面向丰富风格语言-音频预训练的双编码器语音-文本模型
Anuj Diwan, Eunsol Choi, David Harwath
机构
*
Department of Computer Science, The University of Texas at Austin, USA(德克萨斯大学奥斯汀分校计算机科学系)
;
Computer Science and Data Science, New York University, USA(纽约大学计算机科学与数据科学系)
XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
XR-1:通过学习统一的视觉-运动表示实现多功能的视觉-语言-动作模型
Shichao Fan, Kun Wu, Zhengping Che, Xinhua Wang, Di Wu, Fei Liao, Ning Liu, Yixue Zhang, Zhen Zhao, Zhiyuan Xu, Meng Li, Qingjie Liu, Shanghang Zhang, Min Wan, Jian Tang
机构
*
Beijing Innovation Center of Humanoid Robotics, Beijing, China(北京人形机器人创新中心,北京,中国)
;
School of Mechanical Engineering and Automation, Beihang University, Beijing, China(北京航空航天大学机械工程及自动化学院,北京,中国)
;
State Key Laboratory of Virtual Reality Technology and Systems, SCSE, Beihang University, Beijing, China(虚拟现实技术与系统国家重点实验室,SCSE,北京航空航天大学,北京,中国)
;
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University, Beijing, China(多媒体信息处理国家重点实验室,计算机科学学院,北京大学,北京,中国)
Franziska Weindel, Michael Girsch, Reinhard Heckel
机构
*
School of Computation, Information and Technology, Technical University of Munich(计算、信息与技术学院,慕尼黑技术大学)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
机构
*
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
Department of Computer Sciences, University of Wisconsin–Madison(威斯康星大学麦迪逊分校计算机科学系)
;
Department of Statistics, University of Wisconsin–Madison(威斯康星大学麦迪逊分校统计学系)
专题命中
预训练与数据
:large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
机构
*
CSIR-Central Scientific Instruments Organisation, India(印度CSIR-中央科学仪器组织)
;
Academy of Scientific and Innovative Research (AcSIR), Ghaziabad, U.P, India(印度科学与创新研究院(AcSIR))
;
School of Mathematical and Computational Sciences, Massey University, New Zealand(梅西大学数学与计算科学学院)
专题命中
预训练与数据
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Less Is More: Reducing Token Counts Without Compromising Performance
少即是多:在不影响性能的情况下减少词元数量
Gyeongje Cho, Yeonkyoung So, Sangmin Lee, Jaejin Lee
机构
*
Graduate School of Data Science, Seoul National University(数据科学研究生院,首尔国立大学)
;
Department of Computer Science, Seoul National University(计算机科学系,首尔国立大学)
专题命中
预训练与数据
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI