RobeCzech: Czech RoBERTa, a monolingual contextualized language representation model
Milan Straka OrcID: 0000-0003-3295-5576 Affiliation: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics, Malostranské nám. 25, 118 00 Prague, Czech Republic E-mail Jakub Náplava OrcID: 0000-0003-2259-1377 Jana Straková OrcID: 0000-0003-0075-2408 David Samuel OrcID: 0000-0003-2866-1022
Abstract
We present RobeCzech, a monolingual RoBERTa language representation model trained on Czech data. RoBERTa is a robustly optimized Transformer-based pretraining approach. We show that RobeCzech considerably outperforms equally-sized multilingual and Czech-trained contextualized language representation models, surpasses current state of the art in all five evaluated NLP tasks and reaches state-of-the-art results in four of them. The RobeCzech model is released publicly at https://hdl.handle.net/11234/1-3691 and https://huggingface.co/ufal/robeczech-base.
原文 arXiv:2105.11314;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2105.11314v2