
This data descriptor introduces GL-MedQuAD, the first domain-specific biomedical Question Answering (QA) benchmark for the Galician language. Comprising 2,100 parallel records derived from the MedQuAD corpus—specifically from GHR (N=1,467) and NIHSeniorHealth (N=633)—the dataset addresses the critical scarcity of specialized clinical NLP resources in low-resource linguistic scenarios. The dataset was constructed through a hybrid, multi-stage generation-and-curation pipeline: Answer Translation Module: Initial English-to-Galician translations of the medical answers were executed using SalamandraTA 7B Instruct. Question–Answer Alignment & Synthetic Generation Module: To optimize question-to-answer semantic alignment, Gemma 3 27B Instruct was used to directly generate and align the enhanced questions in both English and Galician. Expert Human Curation: To guarantee domain accuracy and linguistic integrity, human post-editing was conducted in compliance with ISO 18587:2017 standards and validated against official Galician medical references (Diccionario galego de termos médicos and Vocabulario de Medicina). The repository includes a comprehensive documentation file along with three cross-aligned CSV datasets linked by a shared record identifier (original_id): README.md: Contains exhaustive documentation regarding dataset usage, column descriptors and error taxonomy codes. Primary Bilingual Parallel Corpus (GL_MedQuAD_bilingual_translations.csv): Includes raw English QA pairs, machine-translated Galician versions of the answers, ISO-curated Galician translations, synthetically generated questions in both English and Galician with curated Galician variants, and instruction-tuned response formats optimized for LLM fine-tuning and conversational healthcare applications. Translation Quality Assessment Dataset (GL_MedQuAD_translation_evaluation.csv)Combines reference-less automated quality metrics (COMETKiwi, LaBSE) with expert human evaluation scores on a 0–5 scale, complemented by a fine-grained 12-category human post-editing error taxonomy. Source-Text Linguistic Complexity Dataset (GL_MedQuAD_linguistic_complexity.csv)Captures the source-text linguistic complexity of the original English answers across three structured, complementary tiers, spanning human-perceived difficulty ratings provided by domain experts, composite complexity indices derived through Multi-Criteria Decision-Making (MCDM) frameworks—such as Analytic Hierarchy Process (AHP), CRITIC weighting, and Hybrid strategies—and fine-grained individual linguistic metrics.
clinical question-answering, linguistic complexity, multi-criteria decision-making, low-resource languages
clinical question-answering, linguistic complexity, multi-criteria decision-making, low-resource languages
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
