
This repository contains the files to reproduce the results from the paper "Impact of Training Instance Selection on Automated Algorithm Selection Models for Numerical Black-box Optimization". Data Collection In this folder, all files used to generate the raw performance data for the set of algorithms are included, as well as the code used to generate the ELA features. For the performance, the packages 'ioh', 'nevergrad', 'modde' and 'modcma' are essential, while 'pflacco' is used for ELA. For all scripts in this folder, the number of parallel threads and the folders for reading function settings (included as 3 csv files in this folder) and storing data should be set before execution. Note that the full performance data exceeds 50GB, so it is not included in this repository. Instead, the results of processing it (using the aocc_extraction script) are included in the 'auc_MABBOB' folder (spread across multiple csv-files, with a version using a different budget factor included as well). The ELA data is included as 'ELA' and 'ELA_BBOB' for the affine combinations and component functions respectively. Data Processing, Analysis and Visualization The remaining reproducibility files can be found in the Reproducibility folder. Within this folder are several notebooks which handle various steps in the pipeline, starting with preprocessing the data collected in the previous steps. This results in some csv-files, which are also included for convenience. Afterwards, the remaining notebooks deal with correlation analysis, instance selection methods, and all included plots from the paper. To match the environment used during our execution of these scripts. a yml-file (to be used with conda or mamba) is available as well.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
