
Ever since the first Short Message Service (SMS) service was introduced in 1993, its popularity has continued to soar over the years such that SMS communication now constitutes a major segment in the spectrum of telecommunication. The popularity and extensive usage has attracted the interest of many researchers to the inherent potential in harvesting data and metadata from collection of SMS corpus for the performance of linguistic, diachronic, normalization and sociolinguistic studies and also in the validation and comparison of different classifiers in SMS spam filters. However, freely available dataset where this type of information can be found for research purposes are quite difficult to obtain. This is mostly due to the confidentiality of SMS where users want to reveal as little of the contents of their phones as possible. This work examines the techniques adopted in the creation of SMS corpus and the ethical consideration involved in the protection of users’ interest and privacy. A critical review of existing work in the field was done to ascertain ethical observations adopted and it was discovered that in other to achieve successful SMS corpus creation, the main consideration is the requirement to protect the rights and interests of the message donors and any other person mentioned in the text messages, without altering the original text in order to gather sufficient metadata information. Participant consent, data anonymization, and ensuring participants’ safe information storage are basic ethical consideration adopted to ensure a successful SMS corpus creation in this work.
Corpus, Metadata, Linguistic, Normalization, Sociolinguistic, SMS, Spam filter
Corpus, Metadata, Linguistic, Normalization, Sociolinguistic, SMS, Spam filter
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
