Powered by OpenAIRE graph
Found an issue? Give us feedback
ZENODOarrow_drop_down
ZENODO
Conference object . 2026
License: CC BY
Data sources: Datacite
ZENODO
Conference object . 2026
License: CC BY
Data sources: Datacite
ZENODO
Conference object . 2026
License: CC BY
Data sources: Datacite
ZENODO
Conference object . 2026
License: CC BY
Data sources: Datacite
addClaim

Clinical Dataset Structure: A Universal Framework for Structuring Clinical Research Datasets

Authors: Gasimova, Aydan; Soundarajan, Sanjay; Gim, Nayoon; Shaffer, Jamie; Owen, Julia; Lee, Aaron; Patel, Bhavesh;

Clinical Dataset Structure: A Universal Framework for Structuring Clinical Research Datasets

Abstract

Clinical studies collect rich multimodal data, such as surveys, wearables, and eye images. There is currently no consensus on how to structure that data into a consistently organized dataset. As a result, clinical datasets are hard to reuse, and datasets from different studies are not readily interoperable. To address this challenge in the AI Ready and Exploratory Atlas for Diabetes Insights (AI-READI) project, we developed the Clinical Dataset Structure (CDS). The CDS is an open-source (CC-BY-4.0), standardized framework for organizing multimodal clinical research data and metadata developed by unifying existing domain-specific standard structures (e.g., Brain Imaging Data Structure), metadata schemas (DataCite, ClinicalTrials.gov), and AI/ML-focused standards (Datasheet, Healthsheet). The CDS recommends organizing multimodal clinical research datasets into one folder per data type, named as per a set convention, where applicable standard structure must then be followed within each folder. It also recommends including multiple metadata files at the root-level in human-friendly (e.g., README, Healthsheet) and machine-friendly (e.g., dataset_description.json, study_description.json) formats that collectively capture over 80 metadata elements. The CDS specification is documented in detailed documentation. We are also building tooling to lower the barrier to adoption. The AI-READI dataset, which is fully structured following the CDS, has already been downloaded over 950 times and contributed to multiple publications, demonstrating real-world impact. In this presentation, we will provide details about the development of the CDS, present its implementation in the AI-READI dataset, and explain how the community can adopt and build on it.

Each year, a rapidly growing volume of datasets is shared across the research community. In clinical studies, many data types are collected per participant, such as surveys, vital signs, eye images, and more. There is currently no consensus on how to organize multimodal data and include related information known as metadata. Standards like BIDS (Brain Imaging Data Structure) exist but only for structuring individual modalities. To address this challenge in the AI Ready and Exploratory Atlas for Diabetes Insights (AI-READI) project, we developed the Clinical Dataset Structure (CDS). The CDS provides a simple and intuitive way to organize clinical research data and metadata in line with the FAIR Principles. It recommends organizing multimodal clinical research datasets into one folder per data type, named as per a set convention, where applicable standard structure must then be followed within each folder.

Keywords

FAIR Principles, Metadata, Clinical Dataset Structure, AI-readiness, Interoperability, Clinical Research Data, Standard, FAIR

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Upload OA version
Are you the author of this publication? Upload your Open Access version to Zenodo!
It’s fast and easy, just two clicks!