Journal refIn Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 5: Industry Track), pages 927-936, Rabat, Morocco, March 2026
Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models
为预训练模型教学旧分词器新词汇:高效的分词器适应方法
Taido Purason, Pavel Chizhov, Ivan P. Yamshchikov, Mark Fishel
机构
*
Institute of Computer Science, University of Tartu(塔尔图大学计算机科学研究所)
;
CAIRO, Technical University of Applied Sciences Würzburg-Schweinfurt(魏玛-施维林应用技术大学)
Vocabulary shapes cross-lingual variation of word-order learnability in language models
词汇如何塑造语言模型中词序可学习性的跨语言变异
Jonas Mayer Martins, Jaap Jumelet, Viola Priesemann, Lisa Beinborn
机构
*
University of Göttingen, Germany(德国哥廷根大学)
;
University of Groningen, Netherlands(荷兰格罗宁根大学)
;
MPI for Dynamics and Self-Organization, Germany(德国动态与自组织研究所)
HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning
HISR: 基于 hindsight 信息的分段过程奖励用于多轮代理强化学习
Zhicong Lu, Zichuan Lin, Wei Jia, Changyuan Tian, Deheng Ye, Peiguang Li, Li Jin, Nayu Liu, Guangluan Xu, Wei Feng
机构
*
Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空航天信息研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Tencent Hunyuan(腾讯文脉)
;
School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院)
Enhancing Lexicon-Based Text Embeddings with Large Language Models
通过大语言模型增强基于词典的文本嵌入
Yibin Lei, Tao Shen, Yu Cao, Andrew Yates
机构
*
University of Amsterdam(阿姆斯特丹大学)
;
University of Technology Sydney(技术悉尼大学)
;
Tencent IEG(腾讯IEG)
;
Johns Hopkins University, HLTCOE(约翰霍普金斯大学,HLTCOE)
Integrating Personality into Digital Humans: A Review of LLM-Driven Approaches for Virtual Reality
将人格融入数字人类:一种基于大语言模型的虚拟现实方法综述
Iago Alves Brito, Julia Soares Dollis, Fernanda Bufon Färber, Pedro Schindler Freire Brasil Ribeiro, Rafael Teixeira Sousa, Arlindo Rodrigues Galvão Filho
Rethinking the Relationship between the Power Law and Hierarchical Structures
重新思考幂律与分层结构之间的关系
Kai Nakaishi, Ryo Yoshida, Kohei Kajikawa, Koji Hukushima, Yohei Oseki
机构
*
RIKEN(理化学研究所)
;
National Institute for Japanese Language and Linguistics(日本语言学研究所)
;
The University of Tokyo(东京大学)
;
Georgetown University(乔治城大学)
;
National Institute of Informatics(信息处理研究所)
Journal refWorkshop on Multilingual and Multicultural Evaluation (MME) of the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL), Pinzhen Chen; Vil{é}m Zouhar; Hanxu Hu; Simran Khanuja; Wenhao Zhu; Barry Haddow; Alexandra Birch; Alham Fikri Aji; Rico Sennrich; Sara Hooker, Mar 2026, Rabbat, Morocco
Are you sure? Measuring models bias in content moderation through uncertainty
你确定吗?通过不确定性测量内容审核中的模型偏差
Alessandra Urbinati, Mirko Lai, Simona Frenda, Marco Antonio Stranisci
机构
*
Laboratory for the Modeling of Biological and Socio-technical Systems, Northeastern University(生物与社会技术系统建模实验室,东北大学)
;
Heriot-Watt University(赫瑞-瓦特大学)
;
aequa-tech(aequa-tech公司)
;
Università del Piemonte Orientale(皮埃蒙特东方大学)
;
Università degli Studi di Torino(托里尼大学)
机构
*
School of Computer Science, Guangdong University of Technology(广东技术大学计算机科学学院)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室)
;
Peng Cheng Laboratory(鹏城实验室)
;
College of Science, Shantou University(汕头大学理学院)
CommentsAccepted at the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL 2025), Long Paper, 19 pages
Journal refProceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 10690-10708. Association for Computational Linguistics, 2025
La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America
La Leaderboard:一种用于西班牙及拉丁美洲各种语言和方言的大型语言模型排行榜
María Grandury, Javier Aula-Blasco, Júlia Falcão, Clémentine Fourrier, Miguel González, Gonzalo Martínez, Gonzalo Santamaría, Rodrigo Agerri, Nuria Aldama, Luis Chiruzzo, Javier Conde, Helena Gómez, Marta Guerrero, Guido Ivetta, Natalia López, Flor Miriam Plaza-del-Arco, María Teresa Martín-Valdivia, Helena Montoro, Carmen Muñoz, Pedro Reviriego, Leire Rosado, Alejandro Vaca, María Estrella Vallecillo-Rodríguez, Jorge Vallego, Irune Zubiaga
机构
*
SomosNLP
;
ETSIT, Universidad Politécnica de Madrid(ETSIT,西班牙马德里理工大学)
;
Barcelona Supercomputing Center(巴塞罗那超级计算中心)
;
Hugging Face
;
Universidad Carlos III de Madrid(马德里卡洛斯三世大学)
;
Instituto de Ingeniería del Conocimiento(知识工程研究所)
;
Centro HiTZ - Ixa, Universidad del País Vasco UPV/EHU(HiTZ-Ixa中心,巴斯克大学)
;
LIACS, Leiden University(LIACS,莱顿大学)
;
Universidad de Jaén(杰嫩大学)
;
Universidad Nacional de Córdoba(科尔多瓦国立大学)
;
Universidad Nacional Autónoma de México(墨西哥国立自治大学)
;
Universidad de la República, Uruguay(乌拉圭共和国大学)
AI总结
La Leaderboard是一个开源项目,旨在评估西班牙及拉丁美洲各种语言和方言的生成式大型语言模型,通过社区驱动的方式促进多样化语言模型的发展。
Journal refProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 32424-32444, July 2025
LLM-as-an-Annotator: Training Lightweight Models with LLM-Annotated Examples for Aspect Sentiment Tuple Prediction
LLM-as-Annotator: 利用LLM标注示例训练轻量模型进行方面情感元组预测
Nils Constantin Hellwig, Jakob Fehle, Udo Kruschwitz, Christian Wolff
机构
*
Media Informatics Group, University of Regensburg, Regensburg, Germany(莱茵河畔大学媒体信息学组)
;
Information Science Group, University of Regensburg, Regensburg, Germany(莱茵河畔大学信息科学组)
AnnoABSA: A Web-Based Annotation Tool for Aspect-Based Sentiment Analysis with Retrieval-Augmented Suggestions
AnnoABSA:一个支持基于方面的情感分析的网页标注工具,具有检索增强的建议
Nils Constantin Hellwig, Jakob Fehle, Udo Kruschwitz, Christian Wolff
机构
*
Media Informatics Group, University of Regensburg, Regensburg, Germany(里根斯堡大学媒体信息学组)
;
Information Science Group, University of Regensburg, Regensburg, Germany(里根斯堡大学信息科学组)
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Peking University Third Hospital(北京大学第三医院)
;
The Hong Kong Polytechnic University(香港理工大学)
;
The University of Hong Kong(香港大学)