Downloads provided by UsageCounts
GenoML is an open source, automated machine learning framework for genomics. GenoML trains and tests multiple algorithms to determine the best fitting solution to a given dataset. Using GenoML, we fit models to every ICD-10 code with at least 1,000 coded samples in the UK Biobank. Observations with the code of interest for each model were considered cases and all remaining observations without the code of interest were considered controls. Cases and controls were randomly sampled at a 1:1 ratio and codes with more than 2,500 associated samples were down sampled to 2,500 observations. Along with normalized age, sex, 5 principal components, and Townsend index score, 31,407 variants were used to build each model. Genotypes were encoded as having 0, 1, or 2 copies of the minor allele. Some codes are specifically male or female, e.g. the codes associated with childbirth are considered female specific. Codes represented by less than 10% of either gender were trained specifically on samples of the majority sex. The resulting models are being made available to other researchers along with a list of SNP names and alleles to facilitate their use on other datasets and populations. These models are trained and tested with European ancestry samples. ICD-10 codes for these models are subcategory level, meaning one level more specific than the main disease category coding (or in some cases, two levels). File name structure follows the schema of ICD-10 subcategory code first, e.g. the code A00.1 is represented as A001. Following the ICD-10 code is the sex code, specifying model training as being on both male and female samples, only male samples, or only female samples. The file “SNP_list.txt” contains the the rs IDs of the SNPs used in each model to allow for model use and replication. Link to code: https://github.com/hll4ce/UKB_GenoML/blob/master/UKB_GenoML_train_discrete.ipynb See also: GenoML ICD-10 code model files (European ancestry, main category codes)
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 1 | |
| downloads | 1 |

Views provided by UsageCounts
Downloads provided by UsageCounts