← Papers

Paper record

Yield prediction through integration of genetic, environment, and management data through deep learning.

Kick DR, Wallace JG, Schnable JC, Kolkman JM, Alaca B, Beissinger TM, Edwards J, Ertl D, Flint-Garcia S, Gage JL, Hirsch CN, Knoll JE, de Leon N, Lima DC, Moreta DE, Singh MP, Thompson A, Weldekidan T, Washburn JD.

G3 (Bethesda, Md.) · 1 Apr 2023 · 10.1093/g3journal/jkad006

Abstract

Accurate prediction of the phenotypic outcomes produced by different combinations of genotypes, environments, and management interventions remains a key goal in biology with direct applications to agriculture, research, and conservation. The past decades have seen an expansion of new methods applied toward this goal. Here we predict maize yield using deep neural networks, compare the efficacy of 2 model development methods, and contextualize model performance using conventional linear and machine learning models. We examine the usefulness of incorporating interactions between disparate data types. We find deep learning and best linear unbiased predictor (BLUP) models with interactions had the best overall performance. BLUP models achieved the lowest average error, but deep learning models performed more consistently with similar average error. Optimizing deep neural network submodules for each data type improved model performance relative to optimizing the whole model for all data types at once. Examining the effect of interactions in the best-performing model revealed that including interactions altered the model's sensitivity to weather and management features, including a reduction of the importance scores for timepoints expected to have a limited physiological basis for influencing yield-those at the extreme end of the season, nearly 200 days post planting. Based on these results, deep learning provides a promising avenue for the phenotypic prediction of complex traits in complex environments and a potential mechanism to better understand the influence of environmental and genetic factors.

Code and data availability

The paper's data availability statement names public Zenodo deposits for the authors' custom python processing scripts (zenodo.org/record/7401113, DOI 10.5281/zenodo.7401113) and the genomic data version (zenodo.org/record/6916775), plus G2F phenotype/weather/soil data via CyVerse DOIs. These are paper-specific, public

No evidence-backed public reproduction asset is currently recorded.

Other versions of this study

Preprints and published versions