Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Dataset . 2025
License: CC BY
Data sources: ZENODO
ZENODO
Dataset . 2025
License: CC BY
Data sources: Datacite
ZENODO
Dataset . 2025
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

ParlMix-UA-RU - Ukrainian Parliamentary Code-Mixing Dataset

Authors: Kanishcheva, Olha; Shvedova, Maria; Dyka, Liudmyla; Husenko, Kristina;

ParlMix-UA-RU - Ukrainian Parliamentary Code-Mixing Dataset

Abstract

ParlMix-UA-RU: Ukrainian Parliamentary Code-Mixing Dataset Overview ParlMix-UA-RU is a specialized, manually annotated linguistic resource designed to facilitate research on code-mixing (CM) and language identification (LID) between Ukrainian and Russian. The dataset is derived from official transcripts of Ukrainian parliamentary sessions (Verkhovna Rada), capturing language dynamics in a formal, high-register political discourse. Dataset Composition & Selection The corpus was developed through a multi-stage process to ensure its utility for complex Natural Language Processing (NLP) tasks: Code-Mixed Sentences: Initially, sentences were extracted based on a threshold of more than two out-of-vocabulary (OOV) tokens relative to monolingual Ukrainian dictionary. This selection typically indicates intra-sentential code-switching or mixed language use. Monolingual Balancing: To improve the robustness of Language Identification models and prevent classification bias, the dataset was subsequently enriched with monolingual Russian sentences from the parlamentary scripts. This ensures a balanced distribution of tokens across the target languages. Annotation Methodology The dataset follows a rigorous "human-in-the-loop" approach to establish a Gold Standard for linguistic research: Manual Tagging: Every token has been manually labeled with its specific language tag by expert linguists. Zero AI Involvement: No automated labeling or AI systems were used for the final annotations, ensuring maximum precision and the preservation of subtle linguistic nuances. Volume: The final dataset contains approximately 150,000 tokens with token-level manual annotations. Research Applications ParlMix-UA-RU is a valuable resource for: Training and benchmarking Language Identification (LID) systems, especially for closely related languages. Computational analysis of intra-sentential code-mixing patterns. Sociolinguistic studies of language use in official governmental settings. Technical Details Language(s): Ukrainian (UA), Russian (RU) Source: Verkhovna Rada transcripts Format: json Annotation level: Token-level language identification Citation & Attribution If you utilize this resource in your academic work, please use the following citation: Kanishcheva, O., Shvedova, M., Dyka, L., & Husenko, K. (2026). ParlMix-UA-RU: Ukrainian Parliamentary Code-Mixing Dataset (Version 1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.14724542

Keywords

Parliament, Language Identification, Ukrainian, code-mixing, Russian, Token-level annotation, code-switching, LID, NLP

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average