
handle: 10394/43009
The main objective of the thesis is to investigate and propose objective text readability measures for Sesotho. This investigation serves as a case study for the development of readability measures for low-resource languages. This thesis assumes that Sesotho texts can be ranked according to their levels of text readability. To measure text readability objectively, we would like to use computational analyses for this. Unfortunately, Sesotho is a low-resource language, lacking sufficient basic language resources necessary for many computational natural language processing methods. We first survey existing readability measures for both high-resource languages, like English, and approaches to develop such measures for low-resource languages. Based on these findings, we identify existing resources needed to develop the measures for Sesotho. As such, various sets of data are collected, including a dictionary lexicon for extracting syllable information, examination texts in all official languages of South Africa from 2008 to 2020, Sesotho Bible translations, Autshumato machine translation texts, and the National Centre for Human Language Technology machine translation texts. In addition, we develop several language resources, such as a syllable annotated wordlist, a corpus-based list of common words in Sesotho, Google Translate machine translations of grade twelve Sesotho examination texts to English, and both pattern-based and rule-based syllabification systems. To create readability measures for Sesotho, we extract text features from Sesotho high school examination texts. The readability values of these texts are computed from their machine-translated English version using existing English readability measures. This allows us to create linear regression models that fit the features against the readability values which leads to readability measures specifically for Sesotho. The structure of the linear regression models for the Sesotho readability measures follows that of the English readability measures. This results in a collection of nine readability measures for Sesotho. Overall, our findings indicate that Sesotho text can be ranked according to levels of readability. Equally important, our results demonstrate that traditional (English) formulas can be adapted to low-resource languages like Sesotho. However, additional efforts may be needed to prepare resources for such projects. We provide a practically useful approach by transitioning from an initial absence of a readability corpus, common vocabulary, and syllable identification to attaining a finalised product. To the best of our knowledge, the findings of this thesis present the first formulas for an indigenous language of South Africa and the first for the Sotho-Tswana language group in Southern Africa. Although our models are trained on educational texts, specifically reading comprehension and summary writing texts, the availability of standardised and objective readability formulas provides a solid starting point for employing machine learning approaches to measure text readability in Sesotho. Furthermore, the methods outlined in the thesis can be employed in the development of readability measures for other low resource languages.
-National Research Foundation (NRF) South Africa -Global Minds Fund Short Research Stays program -Nelson Mandela University through the Teaching Relief Grant -North-West University through Postgraduate Funding
Doctor of Philosophy in Linguistics and Literary Theory, North-West University, Potchefstroom Campus
Doctor of Philosophy (Ph.D.)
Basic Language Resource Kits, Sesotho linguistics, Syllables, Readability, Exam texts, Readability measures, Corpora, Reading ability, Grade 12, BLARK, Syllabification, Sesotho, Traditional readability measures, Text readability
Basic Language Resource Kits, Sesotho linguistics, Syllables, Readability, Exam texts, Readability measures, Corpora, Reading ability, Grade 12, BLARK, Syllabification, Sesotho, Traditional readability measures, Text readability
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
