The application of monolingual and multilingual BERT-based models for text automation tasks in Ukrainian
DOI:
https://doi.org/10.17721/1812-5409.2026/1.26Keywords:
natural language processing, large language models, monolingual and multilingual models, BERTAbstract
The object of this study is mono- and multilingual models based on the BERT architecture. The research focuses on comparing their performance in natural language processing (NLP) tasks, with an emphasis on their application to the Ukrainian language. The methodological framework follows standard approaches to model training and evaluation, relying on publicly available information sources. Experiments were conducted solely on open-source or self-fine-tuned models and covered all existing Ukrainian benchmarks.
Overall, the findings show that both monolingual and multilingual BERT models can be effective for NLP tasks, depending on the target language, specific task, and available resources. While monolingual models often outperform multilingual ones for tasks in their own language, multilingual models can prove advantageous when resources for training language-specific models are limited.
The comparison demonstrates that models trained from scratch on Ukrainian texts–or further fine-tuned on them–generally achieve higher results on Ukrainian NLP tasks than multilingual counterparts. This advantage was evident in named-entity recognition, text classification, and masked-token prediction. Multilingual models, however, achieve competitive or superior performance in "lighter" tasks such as word-sense disambiguation, especially when full training of a monolingual model is not feasible.
These findings underscore the need to adapt or fully train Ukrainian-language models for specialized tasks while still allowing for the cost-effective use of multilingual systems in everyday scenarios. The practical significance of this work lies in providing a consolidated performance map of modern models and laying the groundwork for a full-scale Ukrainian GLUE-like benchmark, which will stimulate further research and improve the accuracy of Ukrainian NLP applications.
Pages of the article in the issue: 195 - 205
Language of the article: Ukrainian
References
Chaplynskyi, D., & Romanyshyn, M. (2024). Introducing NER-UK 2.0: A rich corpus of named entities for Ukrainian. In M. Romanyshyn, N. Romanyshyn, A. Hlybovets, & O. Ignatenko (Eds.), Proceedings of the Third Ukrainian Natural Language Processing Workshop (UNLP) @ LREC-COLING 2024 (pp. 23–29). ELRA and ICCL.
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In J. Burstein, C. Doran, & T. Solorio (Eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol. 1 (Long and Short Papers) (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423
GitHub. (2025). Lang-Uk/ner-uk. https://github.com/lang-uk/ner-uk
Haltiuk, M., & Smywiński-Pohl, A. (2024). LiBERTa: Advancing Ukrainian language modeling through pre-training from scratch (Presentation). In M. Romanyshyn, N. Romanyshyn, A. Hlybovets, & O. Ignatenko (Eds.), Proceedings of the Third Ukrainian Natural Language Processing Workshop (UNLP) @ LREC-COLING 2024 (pp. 120–128). ELRA and ICCL.
Hamotskyi, S., Levbarg, A.-I., & Hänig, C. (2024). Eval-UA-tion 1.0: Benchmark for evaluating Ukrainian (large) language models. In M. Romanyshyn, N. Romanyshyn, A. Hlybovets, & O. Ignatenko (Eds.), Proceedings of the Third Ukrainian Natural Language Processing Workshop (UNLP) @ LREC-COLING 2024 (pp. 109–119). ELRA and ICCL.
Hlybovets, A. M., & Bikchentaiev, M. (2022). Software system for checking plagiarism of Ukrainian texts [Prohramna systema perevirky na plahiat ukrainskykh tekstiv]. Naukovi zapysky NaUKMA. Kompiuterni nauky, 5, 16–25 [in Ukrainian]. https://doi.org/10.18523/2617-3808.2022.5.16-25
Hugging Face. (2024). Uk_ner_web_trf_13class. https://huggingface.co/dchaplinsky/uk_ner_web_trf_13class
Hugging Face. (2025). Roberta-large-wechsel-ukrainian. https://huggingface.co/benjamin/roberta-large-wechsel-ukrainian
Joshi, P., Santy, S., Budhiraja, A., Bali, K., & Choudhury, M. (2020). The state and fate of linguistic diversity and inclusion in the NLP world. In D. Jurafsky, J. Chai, N. Schluter, & J. Tetreault (Eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 6282–6293). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.560
Kiulian, A., Polishko, A., Khandoga, M., Chubych, O., Connor, J., Ravishankar, R., & Shirawalmath, A. (2024). From bytes to borsch: Fine-tuning Gemma and Mistral for the Ukrainian language representation. arXiv preprint arXiv:2404.09138. https://doi.org/10.48550/arXiv.2404.09138
Laba, Y., Mudryi, V., Chaplynskyi, D., Romanyshyn, M., & Dobosevych, O. (2023). Contextual embeddings for Ukrainian: A large language model approach to word sense disambiguation. In M. Romanyshyn (Eds.), Proceedings of the Second Ukrainian Natural Language Processing Workshop (UNLP) (pp. 11–19). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.unlp-1.2
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692. https://doi.org/10.48550/arXiv.1907.11692
Nivre, J., Zeman, D., Ginter, F., & Tyers, F. (2017). Universal dependencies. In A. Klementiev & L. Specia (Eds.), Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Tutorial Abstracts (pp. 1–8). Association for Computational Linguistics. https://aclanthology.org/E17-5001
Pan, X., Zhang, B., May, J., Nothman, J., Knight, K., & Ji, H. (2017). Cross-lingual name tagging and linking for 282 languages. In R. Barzilay & M.-Y. Kan (Eds.), Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Vol. 1: Long Papers) (pp. 1946–1958). Association for Computational Linguistics. https://doi.org/10.18653/v1/P17-1178
Panchenko, D., Maksymenko, D., Turuta, O., Luzan, M., Tytarenko, S., & Turuta, O. (2022). Ukrainian news corpus as text classification benchmark. In O. Ignatenko, V. Kharchenko, V. Kobets, H. Kravtsov, Y. Tarasich, V. Ermolayev, D. Esteban, V. Yakovyna, & A. Spivakovsky (Eds.), ICTERI 2021 Workshops. ICTERI 2021. Communications in Computer and Information Science (Vol. 1635, pp. 550–559). Springer, Cham. https://doi.org/10.1007/978-3-031-14841-5_37
Pires, T., Schlinger, E., & Garrette, D. (2019). How multilingual is multilingual BERT? In A. Korhonen, D. Traum, & L. Màrquez (Eds.), Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 4996–5001). Association for Computational Linguistics. https://doi.org/10.18653/v1/P19-1493
Piskorski, J., Marcińczuk, M., & Yangarber, R. (2024). Cross-lingual named entity corpus for Slavic languages. In N. Calzolari, M.-Y. Kan, V. Hoste, A. Lenci, S. Sakti, & N. Xue (Eds.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 4143–4157). ELRA and ICCL.
Radchenko, V. (2020). We trained the Ukrainian language model. Youscan. https://youscan.io/blog/ukrainian-language-model/
Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. arXiv:1908.10084. https://doi.org/10.48550/arXiv.1908.10084
Ruder, S. (2022). The state of multilingual AI. ruder.io. https://www.ruder.io/stateof-multilingual-ai/
Rybak, P., Mroczkowski, R., Tracz, J., & Gawlik, I. (2020). KLEJ: Comprehensive benchmark for Polish language understanding. In D. Jurafsky, J. Chai, N. Schluter, & J. Tetreault (Eds.), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 1191–1201). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.111
Sanh, V., Wolf, T., Ly, J., & Rush, A. M. (2019). DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108. https://api.semanticscholar.org/CorpusID:203626972
Shaham, U., Aharoni, R., Rappoport, A., & Goldberg, Y. (2024). Multilingual instruction tuning with just a pinch of multilinguality. arXiv preprint arXiv:2401.01854. https://doi.org/10.48550/arXiv.2401.01854
Sido, J., Švec, J., Pecina, P., & Konopík, M. (2021). Czert: Czech BERT-like model for language representation. arXiv preprint arXiv:2103.13031. https://doi.org/10.48550/arXiv.2103.13031
Torge, S., Politov, A., Lehmann, C., Saffar, B., & Tao, Z. (2023). Named entity recognition for low-resource languages – Profiting from language families. Proceedings of the 9th Workshop on Slavic Natural Language Processing 2023 (SlavicNLP 2023) (pp. 1–10), Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.bsnlp-1.1
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. arXiv:1706.03762. https://doi.org/10.48550/arXiv.1706.03762
Virtanen, A., Kanerva, J., Ilo, R., Luoma, J., Luotolahti, J., Salakoski, T., Ginter, F., & Pyysalo, S. (2019). Multilingual is not enough: BERT for Finnish. arXiv preprint arXiv:1912.07076. https://doi.org/10.48550/arXiv.1912.07076
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., & Bowman, S. R. (2019, May). GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 7th International Conference on Learning Representations (ICLR 2019) (poster session) (pp. 353–355). Association for Computational Linguistics. https://doi.org/10.18653/v1/W18-5446
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., & Bowman, S. R. (2019). SuperGLUE: A stickier benchmark for general-purpose language understanding systems. arXiv preprint arXiv:1905.00537. https://doi.org/10.48550/arXiv.1905.00537
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Mykola Glybovets, Danylo Vanin

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
