
## Sesotho Sample Dataset - Next Voices-ZA (South Africa) - Multilingual Speech DatasetThis dataset includes **scripted and unscripted speech** across various domains such as agriculture, health, finance, sports, transport, culture, society and general topics. It is primarily designed for automatic speech recognition (ASR). ### Use Restriction: The persons whose voices are included in this dataset, and the creators and owners of this dataset* do not give consent in any manner or form to, and strictly prohibit any use of this dataset for any form of text-to-speech (TTS), voice cloning, voice synthesis, or any technology or activity intended to replicate, mimic or generate human voices or any technology or activity resulting in the replication, mimicry or generation of human voices. This dataset includes scripted and unscripted speech across various domains such as agriculture, health, finance, sports, transport, culture, society, and general topics. It is primarily designed for use in automatic speech recognition (ASR) tasks. Use of this dataset for any form of text-to-speech (TTS), voice cloning, voice synthesis, or any technology intended to replicate or generate human voices is strictly prohibited. These restrictions are in place until further notice. ## Folder structure The dataset is organised hierarchically as follows: ## Folder StructureANV-ZA-SOT-1h/├── sot/ # Folder for Sesotho│ ├── recorder_uuid/ # Contains all audio files│ │ ├── recording-1731053452.wav│ │ ├── ...│ ├── transcripts.csv # Contains transcripts of all audio recordings│ ├── meta.csv # Contains additional metadata├── README.md # Description of the dataset ## Data Details ### Audio- Format: **16-bit PCM WAV**- Sample rate: **48kHz** ### Transcriptions- Provided in `transcript.csv` with fields: - `file_name`: Name of the audio file. - `transcript`: Text transcription of the audio. - `duration`: Duration of the recording in seconds. - `type`: Scripted or unscripted. ### Metadata- Provided in `meta.csv` with fields: - `recorder_uuid`: Unique speaker identifier. - `age_range`, - `gender` ## Contact PersonPlease contact vukosi.marivate@cs.up.ac.za if you have any questions## CitationTBA ## FundingFunding for this project was generously made possible through a grant from the Bill & Melinda Gates Foundation and a gift from Meta.
Automatic Speech Recognition
Automatic Speech Recognition
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
