Downloads provided by UsageCounts
Cloud masking is a critical step in pre-processing of optical images, for land, water and atmosphere applications. Traditional cloud masking has been done using explicit radiometric tests exploiting physical features such as brightness, whiteness, temperature or specific absorption lines [Ackermann et al, 2010], or spatial features (in/homogeneity) [Eumetsat, 2015], and more or less sophisticated combination of these. Machine learning (ML) is also used since a long time to train a classifier for the purpose of predicting the presence of clouds [Camps-Valls et al 2004]. In recent years ML methods in combination with the now available computation power and easy access to large EO datasets on platforms such as the DIAS’ses, Google Earth Engine or Sentinel Hub, have improved ML based cloud masking significantly [e.g. Zupanc, 2019]. In 2020/2021 under the umbrella of CEOS, ESA and NASA have conducted the Cloud Masking Intercomparison eXercise (CMIX) [CEOS 2019]. One of the major results of this exercise is that “shortcomings and limitations in used reference datasets had been identified“ because each reference datasets only covers a certain aspect of cloud masking. In the context of CMIX, the reference datasets refer to validation datasets. However, this is equally relevant for the construction of training datasets: How is a cloud defined? How are training data collected? Which ranges of other environmental parameters are covered? Furthermore, the CMIX has shown that publicly available datasets have the issue of often being used for training and validation at the same time or being designed for training and then being used for validation, and vice versa. The requirements for ML training data go even further and differ in some respect from requirements for validation data. For example, an ML algorithm may have the requirement to retrieve different cloud classes (cirrus, optically thick cloud, optically thin clouds, …) with the same uncertainty which implies equal distribution of these classes in the training datasets, while validation may require representativity of natural frequency distribution. Eventually, any EO algorithm is an inversion of the radiative transfer (RT) equation, by whatever means. In RT the environmental conditions (such as surface albedo, surface height, cloud, and aerosol optical properties) uniquely determine the radiance field. However, the inversion can have multiple (valid) solutions, it is ambiguous. ML techniques cannot cope with this fact without “helping” them: they would select – more or less randomly – one of the valid solutions. Thus, it is important to either restrict the training dataset and avoid ambiguities, or to include constraints in the ML procedure to add additional knowledge into the training process. In this presentation we will elaborate (a) the importance of training data for the success of the ML method and (b) give examples of restrictions, ambiguities, and possible ways to overcome this issue. We will discuss examples of manually selected training data which is used in our PixBox methodology and conclude with a set of general requirements to be respected when collecting a training dataset for ML for the purpose of cloud masking. Ackerman, Frey, Strabala, Liu, Gumley, Baum, and Menzel, 2010: Discriminating Clear-Sky from Cloud with MODIS - Algorithm Theoretical Basis Document Eumetsat, 2015: MTG-FCI: ATBD for Cloud Mask and Cloud Analysis Product. https://www.eumetsat.int/media/37764 Camps-Valls, Gómez-Chova, Calpe-Maravilla, J. D. Martín-Guerrero, E. Soria-Olivas, L. Alonso-Chordá, José Moreno, 2004: "Robust Support Vector Method for Hyperspectral Data Classification and Knowledge Discovery". IEEE Transactions on Geoscience and Remote Sensing, Vol.42, Issue 7, pp. 1530-1542. CEOS, 2019: CEOS-WGCV ACIX II CMIX Atmospheric Correction Inter-comparison Exercise Cloud Masking Inter-comparison Exercise; https://earth.esa.int/eogateway/events/ceos-wgcv-acix-ii-cmix-atmospheric-correction-inter-comparison-exercise-cloud-masking-inter-comparison-exercise-2nd-workshop?text=cmix Zupanc, 2019: Improving Cloud Detection with Machine Learning. https://medium.com/sentinel-hub/improving-cloud-detection-with-machine-learning-c09dc5d7cf13
Training data, Artificial Intelligence, AI, Machine learning, Earth Observation, Cloud masking, ML
Training data, Artificial Intelligence, AI, Machine learning, Earth Observation, Cloud masking, ML
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 2 | |
| downloads | 1 |

Views provided by UsageCounts
Downloads provided by UsageCounts