Polish ModernBERT:波兰语理解的长短模型
Polish ModernBERT: The Long and Short of Polish Language Understanding
浏览论文内容
中文总结 AI 辅助
该研究推出了Polish ModernBERT系列波兰语编码器,在30项任务中整体性能最优,长上下文任务表现优异且效率更高,在特定波兰语检索基准中也取得最佳结果。
中文摘要 AI 辅助
仅编码器的Transformer在判别式和表示学习任务中仍然有效,但波兰语编码器在很大程度上仍依赖BERT/RoBERTa风格的架构。我们推出了Polish ModernBERT,这是一个包含四个波兰语编码器的系列,有Base和Large两种规模,每种规模都有512-token和8K上下文的变体。我们通过分阶段选择实验调整了ModernBERT的预训练方案,并发布了一个长上下文基准,涵盖法律主题分类、意识形态决策方向预测、文学情节摘要的事实一致性评估以及人权侵犯评估。在30项任务中,Polish ModernBERT在评估的波兰语编码器中实现了最佳整体性能,Base-8K和Large-8K模型分别达到83.99和85.11。在长上下文任务上,8K变体在Base和Large规模下分别将匹配的波兰语RoBERTa-8K基线从67.47提升至77.15,从75.88提升至78.49。Base-8K模型以减少22%的参数(1.49亿对1.9亿)实现了这一提升。在代表性推理设置中的效率测量显示,在512-token和8K设置下,其峰值内存使用和延迟均低于匹配的波兰语RoBERTa基线。Polish ModernBERT-8K-Base还在评估的参数低于3亿的波兰语检索基准中取得了最佳结果。
英文摘要
Encoder-only Transformers remain effective for discriminative and representation-learning tasks, yet Polish encoders still largely rely on BERT/RoBERTa-style architectures. We introduce \textbf{Polish ModernBERT}, a family of four Polish encoders available at Base and Large scales, each with 512-token and 8K context variants. We adapt the ModernBERT pretraining recipe through staged selection experiments and release a long-context benchmark covering legal topic classification, ideological decision-direction prediction, factual-consistency assessment over literary plot summaries, and human-rights violation assessment. Across 30 tasks, Polish ModernBERT achieves the best overall performance among the evaluated Polish encoders, reaching 83.99 and 85.11 for the Base-8K and Large-8K models, respectively. On long-context tasks, the 8K variants improve over matched Polish RoBERTa-8K baselines from 67.47 to 77.15 and from 75.88 to 78.49 at the Base and Large scales, respectively. The Base-8K model achieves this gain with 22\% fewer parameters (149M vs.\ 190M). Efficiency measurements in representative inference setups show lower peak memory usage and latency than matched Polish RoBERTa baselines in both 512-token and 8K settings. Polish ModernBERT-8K-Base additionally achieves the best result on a Polish retrieval benchmark among the evaluated encoders below 300M parameters.
发表机构
- National Information Processing Institute(国家信息处理研究所)
机构由 AI 辅助整理,请以论文原文为准。