
ParlMix-UA-RU: Ukrainian Parliamentary Code-Mixing Dataset Overview ParlMix-UA-RU is a specialized, manually annotated linguistic resource designed to facilitate research on code-mixing (CM) and language identification (LID) between Ukrainian and Russian. The dataset is derived from official transcripts of Ukrainian parliamentary sessions (Verkhovna Rada), capturing language dynamics in a formal, high-register political discourse. Dataset Composition & Selection The corpus was developed through a multi-stage process to ensure its utility for complex Natural Language Processing (NLP) tasks: Code-Mixed Sentences: Initially, sentences were extracted based on a threshold of more than two out-of-vocabulary (OOV) tokens relative to monolingual Ukrainian dictionary. This selection typically indicates intra-sentential code-switching or mixed language use. Monolingual Balancing: To improve the robustness of Language Identification models and prevent classification bias, the dataset was subsequently enriched with monolingual Russian sentences from the parlamentary scripts. This ensures a balanced distribution of tokens across the target languages. Annotation Methodology The dataset follows a rigorous "human-in-the-loop" approach to establish a Gold Standard for linguistic research: Manual Tagging: Every token has been manually labeled with its specific language tag by expert linguists. Zero AI Involvement: No automated labeling or AI systems were used for the final annotations, ensuring maximum precision and the preservation of subtle linguistic nuances. Volume: The final dataset contains approximately 150,000 tokens with token-level manual annotations. Research Applications ParlMix-UA-RU is a valuable resource for: Training and benchmarking Language Identification (LID) systems, especially for closely related languages. Computational analysis of intra-sentential code-mixing patterns. Sociolinguistic studies of language use in official governmental settings. Technical Details Language(s): Ukrainian (UA), Russian (RU) Source: Verkhovna Rada transcripts Format: json Annotation level: Token-level language identification Citation & Attribution If you utilize this resource in your academic work, please use the following citation: Kanishcheva, O., Shvedova, M., Dyka, L., & Husenko, K. (2026). ParlMix-UA-RU: Ukrainian Parliamentary Code-Mixing Dataset (Version 1.0.0) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.14724542
Parliament, Language Identification, Ukrainian, code-mixing, Russian, Token-level annotation, code-switching, LID, NLP
Parliament, Language Identification, Ukrainian, code-mixing, Russian, Token-level annotation, code-switching, LID, NLP
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
