Detailed information on all experimental lines, including their genotypes, phenotypic, and environmental data, is available at https://doi.org/10.60867/00000010 , https://doi.org/10.60867/00000003 , and https://doi.org/10.60867/00000011 , respectively.
Open resource ↗10.60867 · 10.60867/00000010 · lines:31-42Paper record
Machine learning to predict genotypes and genotype-environment interaction associated with complex traits for genomic selection.
Plant phenomics (Washington, D.C.) · 19 May 2026 · 10.1016/j.plaphe.2026.100224
Abstract
Genomic selection (GS) can accelerate crop breeding and enhance selection efficiency. However, accurately predicting genomic estimated breeding values (GEBVs) for complex traits and applying GS in diverse environments remains challenging. To address these issues, we developed a novel hybrid method capable of modelling gene-gene and gene-environment interactions. This method offers precise predictions of phenotypic performance for complex traits, identifies haplotypes associated with desirable phenotypes, and enables prediction of optimal haplotypes tailored to specific environments. We evaluated the approach using a dataset of 855 barley lines, with phenotypic data for grain yield and flowering time collected across multiple environments. The model incorporated 30,543 SNPs, nine soil parameters, and six daily environmental variables, achieving high prediction accuracies, with correlation coefficients of 0.93 for flowering time and 0.82 for grain yield. Our method identified 10 haplotype blocks significantly associated with flowering time and 13 blocks with grain yield, collectively accounting for over 90% of the total genetic variance. Additionally, we predicted the phenotypic effects of each haplotype and identified elite varieties carrying the most favourable haplotypes for crossing design and selection. The method also allows prediction of untested genotype × environment combinations, enabling selection of optimal genotypes for targeted environments. To facilitate its application, we developed a web-based interface (accessible at [https://penghaowang.shinyapps.io/shinygui/]), which enables breeders to identify optimal haplotypes and the varieties that carry them, streamlining the process of haplotype-based, environment-informed breeding. We note that the reverse prediction framework is currently applied on a single-trait basis and does not resolve multi-trait trade-offs such as between flowering time and yield, which remains a topic for future extensions.
Code and data availability
The paper deposits its barley genotype, phenotype, and environmental datasets at three DOI repositories, and its analysis source code on GitHub, plus a public Shiny web tool.
Detailed information on all experimental lines, including their genotypes, phenotypic, and environmental data, is available at https://doi.org/10.60867/00000010 , https://doi.org/10.60867/00000003 , and https://doi.org/10.60867/00000011 , respectively.
Open resource ↗10.60867 · 10.60867/00000003 · lines:31-42Detailed information on all experimental lines, including their genotypes, phenotypic, and environmental data, is available at https://doi.org/10.60867/00000010 , https://doi.org/10.60867/00000003 , and https://doi.org/10.60867/00000011 , respectively.
Open resource ↗10.60867 · 10.60867/00000011 · lines:31-42All the data and source codes have been uploaded to GitHub and can be accessed under the GNU Open License at: https://github.com/pwang2019/GxE_Model .
Open resource ↗github.com/pwang2019/GxE_Model · lines:196-205