
doi: 10.4018/407607
Arabic is a first language for more than 300 million people. It has some unique features that can make it one of the most complex languages, such as multiple derivatives, unlimited vocabulary, diacritics, and others. Preprocessing Arabic text is an essential step in order to prepare text for Natural Language Processing (NLP) purposes. This article provides a comparison study of several preprocessing tools for Arabic text. It explains the challenges in pre-processing the Arabic language as well as the techniques that used in every particular tool. However, the authors used the PRISMA for reporting the systematic reviews, which they started with screening 200 articles and ended-up with including only 30 articles. After reviewing these articles deeply, the results show that different tools such as AMIRA, CAMel ,and NLP packages added value in text-preprocessing. However, most of this papers considered that the ambiguity in Arabic orthography as well as the dialectal variants are the most challenges in Arabic NLP.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
