WSEAS Transactions on Information Science and Applications
Print ISSN: 1790-0832, E-ISSN: 2224-3402
Volume 23, 2026
AraBART-based Arabic Lemmatization
Authors: , ,
Search Articles
Abstract: Arabic, an inflectional language with a rich morphology and complex syntactic structures, demands robust approaches for effective normalization and lemmatization. While, most existing NLP models focus on English, Modern Standard Arabic remains understudied, particularly in lemmatization tasks. In this paper, we introduce AraBART, the first Arabic model to feature an end-to-end pre-trained encoder-decoder, leveraging the BART architecture. We used the Arabic-PADT UD Treebank and Farasa corpus. Performance was assessed with accuracy, precision, recall, and F1 score. The results show that AraBART surpasses strong baselines, including transformer-based models such as AraT5, mT5, and BERT, achieving a 4.72% improvement in lemmatization accuracy. More importantly, AraBART has achieved an accuracy of 94.71%, approaching the performance of Farasa (97.32%). Furthermore, incorporating parts of the Farasa corpus into our training process showed a clear improvement in accuracy, revealing AraBART's effectiveness for broad applications in Arabic NLP tasks, including text summarization and machine translation.
Keywords:
Arabic NLP, Lemmatization, Transformers, Modern Standard Arabic, AraBART, AraT5, MT5, BERT
Pages: 1-11
DOI: 10.37394/23209.2026.23.1