The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese
首个中文BabyLM挑战:训练数据高效且认知合理的中文语言模型
Siyuan Song, Zhiheng Qian, Yunhao Zhang, Linyang He, Xiaozhe Ji, Yingxin Lin, Hongao Zhu, Chongtian Shao, Chuhan Lang, Luan Li, Rui Wang, Renfen Hu, Shaonan Wang, Hai Hu
机构
*
Princeton University(普林斯顿大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Chinese Academy of Sciences(中国科学院)
;
Columbia University(哥伦比亚大学)
;
Beijing Normal University(北京师范大学)
;
Tsinghua University(清华大学)
;
University of California San Diego(加利福尼亚大学圣地亚哥分校)
;
The Hong Kong Polytechnic University(香港理工大学)
TESSERA v2: Scaling Pixel-wise Earth Foundation Models
TESSERA v2:扩展逐像素地球基础模型
Zhengpeng Feng, Sadiq Jaffer, Ira Shokar, Jovana Knezevic, James Ball, Pedro Sousa, Mark Elvers, Madeline Lisaius, Clement Atzberger, Robin Young, Aneesh Naik, Niall Robinson, David Coomes, Anil Madhavapeddy, Srinivasan Keshav
机构
*
University of Cambridge(剑桥大学)
;
NVIDIA(英伟达)
;
dClimate Labs(dClimate实验室)
KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report
KletterMix: 攀登高质量德语预训练数据
Maurice Kraus, Ruben Härle, Sebastian Sztwiertnia, Abbas Goher Khan, Mehdi Ali, Michael Fromm, Nicolas Flores-Herr, Kristian Kersting
机构
*
AI & ML Group, TU Darmstadt(人工智能与机器学习小组,德累斯顿技术大学)
;
Lab1141
;
Lamarr Institute(拉马尔研究所)
;
Fraunhofer IAIS(弗劳恩霍夫人工智能研究所)
;
hessian.AI(海斯坦.AI)
;
German Research Center for AI (DFKI)(德国人工智能研究中心(DFKI))
;
Centre for Cognitive Science, TU Darmstadt(认知科学中心,德累斯顿技术大学)
Detecting LLM-Generated Tokens in Human--LLM Coauthored Text
在人类与大语言模型合作撰写的文本中检测大语言模型生成的令牌
Yangjun Lu, Hongyi Zhou, Fabian Spill, Kai Ye, Chengchun Shi, Jin Zhu
机构
*
School of Mathematics, University of Birmingham(伯明翰大学数学学院)
;
School of Statistics and Data Science, Shanghai University of Finance and Economics(上海财经大学统计与数据科学学院)
;
Department of Statistics, The London School of Economics and Political Science(伦敦政治经济学院统计系)
Perturbation is All You Need for Extrapolating Language Models
扰动是语言模型外推所需的一切
Zetai Cen, Jin Zhu, Xinwei Shen, Chengchun Shi
机构
*
School of Mathematics, University of Bristol(布里斯托大学数学系)
;
School of Mathematics, University of Birmingham(伯明翰大学数学系)
;
Department of Statistics, University of Washington(华盛顿大学统计系)
;
Department of Statistics, London School of Economics and Political Science(伦敦政治经济学院统计系)
专题命中
预训练与数据
:language model(title,abstract);large language model(abstract);分类 cs.LG
Knowledgeless Language Models: Suppressing Parametric Recall for Evidence-Grounded Language Modeling
无知识语言模型:抑制参数化记忆以进行基于证据的语言建模
Roi Cohen, Yvan Carré, Nick Lechtenbörger, Hendrik Droste, Lucas Kerschke, Russa Biswas, Gerard de Melo, Jan Buys
机构
*
HPI / University of Potsdam(波茨坦大学哈索·普拉特纳数字工程学院)
;
African Institute for Mathematical Sciences(非洲数学科学研究所)
;
Department of Computer Science Aalborg University(奥尔堡大学计算机科学系)
;
Department of Computer Science University of Cape Town(开普敦大学计算机科学系)
;
Polytechnique Montréal(蒙特利尔理工大学)