Powered by OpenAIRE graph
Found an issue? Give us feedback
addClaim

Measuring Text Readability in Sesotho

Authors: Sibeko, Johannes;

Measuring Text Readability in Sesotho

Abstract

The main objective of the thesis is to investigate and propose objective text readability measures for Sesotho. This investigation serves as a case study for the development of readability measures for low-resource languages. This thesis assumes that Sesotho texts can be ranked according to their levels of text readability. To measure text readability objectively, we would like to use computational analyses for this. Unfortunately, Sesotho is a low-resource language, lacking sufficient basic language resources necessary for many computational natural language processing methods. We first survey existing readability measures for both high-resource languages, like English, and approaches to develop such measures for low-resource languages. Based on these findings, we identify existing resources needed to develop the measures for Sesotho. As such, various sets of data are collected, including a dictionary lexicon for extracting syllable information, examination texts in all official languages of South Africa from 2008 to 2020, Sesotho Bible translations, Autshumato machine translation texts, and the National Centre for Human Language Technology machine translation texts. In addition, we develop several language resources, such as a syllable annotated wordlist, a corpus-based list of common words in Sesotho, Google Translate machine translations of grade twelve Sesotho examination texts to English, and both pattern-based and rule-based syllabification systems. To create readability measures for Sesotho, we extract text features from Sesotho high school examination texts. The readability values of these texts are computed from their machine-translated English version using existing English readability measures. This allows us to create linear regression models that fit the features against the readability values which leads to readability measures specifically for Sesotho. The structure of the linear regression models for the Sesotho readability measures follows that of the English readability measures. This results in a collection of nine readability measures for Sesotho. Overall, our findings indicate that Sesotho text can be ranked according to levels of readability. Equally important, our results demonstrate that traditional (English) formulas can be adapted to low-resource languages like Sesotho. However, additional efforts may be needed to prepare resources for such projects. We provide a practically useful approach by transitioning from an initial absence of a readability corpus, common vocabulary, and syllable identification to attaining a finalised product. To the best of our knowledge, the findings of this thesis present the first formulas for an indigenous language of South Africa and the first for the Sotho-Tswana language group in Southern Africa. Although our models are trained on educational texts, specifically reading comprehension and summary writing texts, the availability of standardised and objective readability formulas provides a solid starting point for employing machine learning approaches to measure text readability in Sesotho. Furthermore, the methods outlined in the thesis can be employed in the development of readability measures for other low resource languages.

-National Research Foundation (NRF) South Africa -Global Minds Fund Short Research Stays program -Nelson Mandela University through the Teaching Relief Grant -North-West University through Postgraduate Funding

Doctor of Philosophy in Linguistics and Literary Theory, North-West University, Potchefstroom Campus

Doctor of Philosophy (Ph.D.)

Country
South Africa
Related Organizations
Keywords

Basic Language Resource Kits, Sesotho linguistics, Syllables, Readability, Exam texts, Readability measures, Corpora, Reading ability, Grade 12, BLARK, Syllabification, Sesotho, Traditional readability measures, Text readability

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Upload OA version
Are you the author of this publication? Upload your Open Access version to Zenodo!
It’s fast and easy, just two clicks!