
doi: 10.3390/a18120733
High-dimensional analytical datasets, such as those generated by inductively coupled plasma–mass spectrometry (ICP-MS), require robust computational frameworks for dimensionality reduction, classification, and model validation. This study presents a comparative evaluation of Linear Discriminant Analysis (LDA) and Partial Least Squares Discriminant Analysis (PLS-DA) algorithms applied to multivariate chemometric data for food origin authentication. The research employs a workflow that integrates Principal Component Analysis (PCA) for feature extraction, followed by supervised classification using LDA and PLS-DA. Model performance and stability were systematically assessed. The dataset comprised 28 apple samples from four geographical regions and was processed with normalization, scaling, and transformation prior to modeling. Each model was validated via leave-one-out cross-validation and evaluated using accuracy, sensitivity, specificity, balanced accuracy, detection prevalence, p-value, and Cohen’s Kappa. The results demonstrate that, as a linear projection-based classifier, LDA provides higher robustness and interpretability in small and unbalanced datasets. In contrast, PLS-DA, which is optimized for covariance maximization, exhibits higher apparent sensitivity but lower reproducibility under similar conditions. The study also emphasizes the importance of dimensionality reduction strategies, such as PCA-based variable selection versus latent space extraction in PLS-DA, in controlling overfitting and improving model generalizability. The proposed algorithmic workflow provides a reproducible and statistically sound approach for evaluating discriminant methods in chemometric classification.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 6 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Top 10% |
