Downloads provided by UsageCounts
doi: 10.1145/3578245.3584935 , 10.48550/arxiv.2301.12166 , 10.5281/zenodo.7661026 , 10.5281/zenodo.7661027
arXiv: 2301.12166
handle: 11311/1245197 , 11311/1232165
doi: 10.1145/3578245.3584935 , 10.48550/arxiv.2301.12166 , 10.5281/zenodo.7661026 , 10.5281/zenodo.7661027
arXiv: 2301.12166
handle: 11311/1245197 , 11311/1232165
Survival analysis studies time-modeling techniques for an event of interest occurring for a population. Survival analysis found widespread applications in healthcare, engineering, and social sciences. However, the data needed to train survival models are often distributed, incomplete, censored, and confidential. In this context, federated learning can be exploited to tremendously improve the quality of the models trained on distributed data while preserving user privacy. However, federated survival analysis is still in its early development, and there is no common benchmarking dataset to test federated survival models. This work provides a novel technique for constructing realistic heterogeneous datasets by starting from existing non-federated datasets in a reproducible way. Specifically, we propose two dataset-splitting algorithms based on the Dirichlet distribution to assign each data sample to a carefully chosen client: quantity-skewed splitting and label-skewed splitting. Furthermore, these algorithms allow for obtaining different levels of heterogeneity by changing a single hyperparameter. Finally, numerical experiments provide a quantitative evaluation of the heterogeneity level using log-rank tests and a qualitative analysis of the generated splits. The implementation of the proposed methods is publicly available in favor of reproducibility and to encourage common practices to simulate federated environments for survival analysis.
FOS: Computer and information sciences, Computer Science - Machine Learning, federated learning, datasets, survival analysis, Machine Learning (cs.LG)
FOS: Computer and information sciences, Computer Science - Machine Learning, federated learning, datasets, survival analysis, Machine Learning (cs.LG)
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 9 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 10% |
| views | 7 | |
| downloads | 10 |

Views provided by UsageCounts
Downloads provided by UsageCounts