← Papers

Paper record

Disentangling genotype and environment specific latent features for improved trait prediction using a compositional autoencoder.

Powadi A, Jubery TZ, Tross MC, Schnable JC, Ganapathysubramanian B.

Frontiers in plant science · 16 Dec 2024 · 10.3389/fpls.2024.1476070

Abstract

In plant breeding and genetics, predictive models traditionally rely on compact representations of high-dimensional data, often using methods like Principal Component Analysis (PCA) and, more recently, Autoencoders (AE). However, these methods do not separate genotype-specific and environment-specific features, limiting their ability to accurately predict traits influenced by both genetic and environmental factors. We hypothesize that disentangling these representations into genotype-specific and environment-specific components can enhance predictive models. To test this, we developed a compositional autoencoder (CAE) that decomposes high-dimensional data into distinct genotype-specific and environment-specific latent features. Our CAE framework employed a hierarchical architecture within an autoencoder to effectively separate these entangled latent features. Applied to a maize diversity panel dataset, the CAE demonstrated superior modeling of environmental influences and out-performs PCA (principal component analysis), PLSR (Partial Least square regression) and vanilla autoencoders by 7 times for 'Days to Pollen' trait and 10 times improved predictive performance for 'Yield'. By disentangling latent features, the CAE provided a powerful tool for precision breeding and genetic research. This work has significantly enhanced trait prediction models, advancing agricultural and biological sciences.

Code and data availability

The paper's data availability statement explicitly deposits the hyperspectral reflectance dataset and trained model weights on Figshare, and the authors' analysis code on a public Bitbucket repository (baskargroup/cae_hyperspectral). Both are paper-specific, public, and actionable.

Datasetpublic

te the technical advantages of disentanglement, it is not immediately clear how to connect these disentangled features to biological insights. Data availability statement The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found below: https://figshare.com/articles/dataset/Hyperspectral_reflectance_data_molecular_and_weights_for_trained_model/24808491/4 ; https://bitbucket.org/baskargroup/cae_hyperspectral/src/main/ . Author contributions AP: Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft, Writing – review & editing. TJ: Conceptualization,

Open resource ↗figshare · 24808491 · lines:460-495