Downloads provided by UsageCounts
Tunizi is the first 100% Tunisian Arabizi sentiment analysis dataset. Tunisian Arabizi is the representation of the tunisian dialect written in Latin characters and numbers rather than Arabic letters.We gathered comments from social media platforms that express sentiment about popular topics. For this purpose, we extracted 100k comments using public streaming APIs. Tunizi was preprocessed by removing links, emoji symbols, and punctuations. The collected comments were manually annotated using an overall polarity: positive (1), negative (-1) and neutral (0) class. We divided the dataset into separate training, validation and test sets, with a ratio of 7:1:2 with a balanced split where the number of comments from positive class and negative class are almost the same.
This dataset contains comments written in Tunisian Arabizi which represents the Tunisian dialect written in Latin letters and numbers.
Tunisian Arabizi, Dialect
Tunisian Arabizi, Dialect
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 160 | |
| downloads | 22 |

Views provided by UsageCounts
Downloads provided by UsageCounts