Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ HAL Université de To...arrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
HAL Université de Tours
Conference object . 2023
addClaim

This Research product is the result of merged Research products in OpenAIRE.

You have already added 0 works in your ORCID record related to the merged Research product.

Automatic Transcription of Tunisian Spoken Arabic : A bridge to linguistic analysis

Authors: Goncalves Teixeira, Daphne; Bahri, Mohamed; Vancaeyzeele, Charles;

Automatic Transcription of Tunisian Spoken Arabic : A bridge to linguistic analysis

Abstract

This article presents a project for the transcription of the spoken Arabic corpus in Tunis, conducted by doctoral students at the Laboratoire Ligérien de Linguistique.The language, which is poorly documented and lacks a standard orthography, poses significant challenges for manual transcription due to vocalic variations. To mitigate biases, an automated transcription was explored using existing models for other languages. The main aim was to reduce the workload and document the language while avoiding a strict reference to standard Arabic. The methodology involves three stages : automated transcription using Whisper AI, transliteration into Latin characters (Buckwalter standards), and harmonization of divergent characters. Thisapproach facilitates comparison with manual transcriptions and will enable the establishment of an orthographic convention for vowels, as well as the identification of phonological processes in synchrony through the compilation of a phonemic inventory.This project will contribute to a better understanding of spoken Arabic in Tunis.

Cet article expose un projet de transcription du corpus de l’arabe parlé à Tunis, mené par des doctorants du Laboratoire Ligérien de Linguistique. La langue, peu documentée et sans orthographe standard, présente des défis majeurs en transcription manuelle en raison des variations vocaliques. Pour pallier les biais, une transcription automatique a été explorée, utilisant des modèles existants pour d’autres langues.L’objectif était de réduire la charge de travail et de documenter la langue tout en évitant une référence stricte à l’arabe standard. La méthodologie inclut trois étapes : transcription automatique avec Whisper AI, translittération en caractères latins (normes de Buckwalter), et harmonisation des caractères divergents. Cette approche facilite la comparaison avec les transcriptions manuelles et permettra d’établir une convention orthographique pour les voyelles, ainsi que d’identifier les processus phonologiques en synchronie grâce à un répertoire phonémique constitué. Ce projet contribuera à une meilleure compréhension de l’arabe parlé à Tunis.

Country
France
Keywords

Transcription automatique, Distributional analysis, corpus oraux, Analyse distributionnelle, [INFO.INFO-CL] Computer Science [cs]/Computation and Language [cs.CL], Oral corpora, Automatic transcription, [SCCO.LING] Cognitive science/Linguistics

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green