
HER2 Breast Cancer Digital Image Dataset (ADEL Dataset). We have developed the first Kazakhstani dataset of digital images for HER2 breast cancer analysis. The dataset consists of images sourced from the pathological archives of the Department of Pathology at the Almaty Oncology Center and the Kazakh Institute of Oncology. Each image is labeled with HER2 expression levels manually assessed by experienced pathologists, with in situ hybridization (ISH) performed in equivocal cases to establish ground truth. The dataset contains 418 images in PNG format. The annotations can be found in the file her2_dataset/labels.csv. HER2 IHC High-Resolution Dataset (Version 0.2) This is **Version 0.2** of the HER2 Immunohistochemistry (IHC) dataset. It includes **high-resolution `.tar` archives** containing processed image tiles extracted from whole-slide images (WSIs). This dataset is hosted on the [Hugging Face Hub]. Digital images were acquired via a fully automated digital system (KFB PRO 120 scanner) at INVIVO LLP with 40x magnification and one focusing layer, ranging in size from 50 MB to 2 GB, depending on the size of the tissue sample fixed on the original slide. The dataset consists of 418 images, which were preprocessed using a conversion script that transformed SVS files into sub-images with a 1:1 aspect ratio in JPEG format. A non-overlapping sliding window approach was applied to generate these sub-images, optimized for machine learning applications. The compressed .png version of the dataset may serve as a visual reference to the characteristics of the original images. --- π₯ Download Instructions πΈ Option 1: Using Python (via `datasets` library) python from datasets import load_dataset dataset = load_dataset("aidosSarsembayev/adel_dataset_1") This will provide access to metadata or a data loader if defined. For raw files (e.g., `.tar`), use git-lfs: πΈ Option 2: Using `git` + `git-lfs` (recommended for large files) bash git lfs install git clone https://huggingface.co/datasets/aidosSarsembayev/adel_dataset_1 This will download all parts of the dataset, including large `.tar` files. --- π Dataset Contents The dataset consists of multiple `.tar.gz` archive files: - `HER2_001_009.tar.gz`- `HER2_010_019.tar.gz`- ...- `HER2_420_429.tar.gz` Each archive contains high-resolution tiles from several HER2 slides. A JSON manifest (`manifest.json`) is provided, mapping each archive to the slide IDs it contains. --- π§ͺ Usage This dataset is intended for research on: - HER2 status classification- Digital pathology and WSI analysis- IHC image processing --- π§ Processing Scripts To reproduce or analyze the dataset, use the scripts provided in the following repository: π [GitHub β HER2 Data Processing]() --- π Citation and License Please refer to the associated Zenodo record or publication for citation and licensing terms. Creative Commons licenses may apply (e.g., CC-BY 4.0). --- Maintained by: [@asarsembayev](https://huggingface.co/aidosSarsembayev)
HER2, AI in breast cancer diagnosis, Medical imaging dataset, HER2 classification, Immunohistochemistry, In Situ Hybridization
HER2, AI in breast cancer diagnosis, Medical imaging dataset, HER2 classification, Immunohistochemistry, In Situ Hybridization
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
