← Papers

Paper record

Empirically calibrated simulations reveal the limits of phenotypic clustering algorithms for biodiversity assessment in data-scarce crops.

Naino Jika AK.

PloS one · 17 Dec 2025 · 10.1371/journal.pone.0329254

Abstract

Clustering algorithms are widely used for phenotypic characterization and germplasm management, particularly in data-scarce crops such as neglected and underutilized species (NUS) that lack genomic resources. However, their performance under biologically realistic conditions remains poorly understood. Standard clustering methods commonly applied in crop research often assume distinct, isotropic, and homogeneous clusters, assumptions rarely satisfied in real-world phenotypic datasets. We developed a flexible and empirically calibrated simulation framework, using phenotypic data from West African fonio (Digitaria exilis), to benchmark the performance of eleven clustering algorithms under both idealized and realistic scenarios. Our simulations integrated heterogeneous trait distributions (normal, gamma), strong inter-trait correlations (up to r = -0.84), heteroscedasticity, and moderate population structure (mean Pst = 0.16 ± 0.001, achieved through iterative calibration). Each scenario was replicated 100 times, with clustering accuracy evaluated using external (ARI, NMI) and internal (Silhouette, Davies-Bouldin) validation metrics under standardized conditions. The results revealed consistently poor algorithm performance under realistic conditions (e.g., ARI < 0.07), including for widely used methods in Neglected and Underutilized Species (NUS) research such as K-means, GMM, and PAM. Notably, conventional validation metrics failed to detect biologically meaningful structure revealed by geometric diagnostics, highlighting a critical methodological limitation. Performance markedly improved under idealized conditions, validating our simulation framework. These findings highlight the risk of overinterpreting clustering outputs from weakly structured phenotypic datasets and expose key limitations in current biodiversity analysis practices, particularly those guiding plant genetic resource conservation programs. We provide an open-source R-based diagnostic tool, with parameter specifications to assist practitioners in selecting reproducible and interpretable clustering approaches for germplasm management and biodiversity assessment in data-scarce crops.

Code and data availability

The paper's Data Availability statement explicitly deposits the complete R simulation/clustering/evaluation script on Zenodo (DOI 10.5281/zenodo.15877863), a paper-specific, publicly actionable code asset. The empirical fonio trait data belong to a prior cited study (Bio et al.), not this paper, and supporting files/DO

Codepublic

the complete R script used to simulate phenotypic datasets, apply clustering algorithms, and compute evaluation metrics is publicly available on Zenodo: https://doi.org/10.5281/zenodo.15877863

Open resource ↗Zenodo · 10.5281/zenodo.15877863 · lines:107-122