Downloads provided by UsageCounts
{"references": ["Acoustic models of Brazilian Portuguese Speech based on Neural Transformers - Gauy, Marcelo e Finger, Marcelo 2022", "MENDES, R.B. (2013) Projeto SP2010: Amostra da fala paulistana. Dispon\u00edvel em .", "Gon\u00e7alves, S. C. L. Projeto ALIP (Amostra Lingu\u00edstica do Interior Paulista) e banco de dados Iboruna: 10 anos de contribui\u00e7\u00e3o com a descri\u00e7\u00e3o do portugu\u00eas brasileiro. ESTUDOS LINGU\u00cdSTICOS (S\u00c3O PAULO. 1978), v. 48, p. 276-297, 2019.", "RASO, T. ; MELLO, H. . The C-ORAL-BRASIL I: Reference Corpus for Informal Spoken Brazilian Portugues. Lecture Notes on Artificial Intelligence, v. 7243, p. 362-368, 2012.", "Oliviera Jr., M. (2016). NURC Digital Um protocolo para a digitaliza\u00e7\u00e3o, anota\u00e7\u00e3o, arquivamento e dissemina\u00e7\u00e3o do material do Projeto da Norma Urbana Lingu\u00edstica Culta (NURC). CHIMERA: Revista De Corpus De Lenguas Romances Y Estudios Ling\u00fc\u00edsticos, 3(2), 149\u2013174. Recuperado a partir de https://revistas.uam.es/chimera/article/view/6519.", "A linguagem falada culta na cidade de S\u00e3o Paulo: materiais para seu estudo - Castilho, Ataliba Teixeira de e Pretti, Dino 1986", "Acervo Certas Palavras- Cat\\'{a}logo 1981-1996 - Teixeira, Carmem Silva P. 1997"]}
This repository contains all the pretraining datasets used in the paper: Acoustic models of Brazilian Portuguese Speech based on Neural Transformers by Marcelo Gauy and Marcelo Finger. These datasets are part of a collection of datasets from the TaRSila project (see https://sites.google.com/view/tarsila-c4ai). The audios published here were in part also published with annotations and transcriptions as the CORAA dataset (see https://github.com/nilc-nlp/CORAA). Here we publish the original raw audios from the following datasets (without transcriptions) - ALIP, C-Oral, SP2010, NURC-Recife, NURC-São Paulo and Programa Certas Palavras. In total, the datasets contain about 800 hours of Brazilian Portuguese Speech. The audios have been converted to mp3 to facilitate the upload. ALIP, C-Oral and SP2010 are integrally contained in one file each. Programa Certas Palavras and NURC-Recife are split in 3 parts each, while NURC-SP is split in 7 parts of roughly equal size. More information on the datasets can be found in the paper Acoustic models of Brazilian Portuguese Speech based on Neural Transformers as well as on the original references which created these datasets.
Unsupervised pretraining (self supervision), Brazilian Portuguese speech
Unsupervised pretraining (self supervision), Brazilian Portuguese speech
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 46 | |
| downloads | 17 |

Views provided by UsageCounts
Downloads provided by UsageCounts