Paper record
A Synthetic Data Generation Pipeline for Improving the Segmentation of Roots in Micro‐CT Images of Soil
European Journal of Soil Science. · 1 Jan 2025
Abstract
Machine learning (ML) models for image segmentation typically require a significant amount of accurately annotated data for training, which is rarely readily available in plant and soil science datasets due to the high time and monetary costs of manually labelling the images. Training datasets can be augmented with synthetically generated images that aim to match the visual features and biological properties of the original dataset. Segmentation masks can be created automatically during the synthetic image generation process, removing the need for tedious manual annotation and ensuring high accuracy of the labels. We present an adaptable semi‐automatic pipeline for creating annotated synthetic micro‐computed tomography (micro‐CT) volumes at scale using the 3D modelling tool Blender, and we demonstrate our method using a dataset of micro‐CT images of tomato plant roots embedded in sieved soil columns. First, the foreground is generated using a mathematical L‐system model to give a 3D model of the target sample. Then, the surrounding material is created and textured to simulate the relative density of the materials in which the object is embedded. The final stage is to render the images by slicing the volume at defined regular intervals, generating both the synthetic micro‐CT image and the corresponding labels at each slice. We use our synthetically generated images alongside real data to create augmented datasets to train a U‐Net‐based segmentation model. Our results demonstrate that when there is a small amount of real annotated data available, using synthetic data in the training dataset can improve the segmentation accuracy, and we show the impact of varying the texturing process.
Code and data availability
公開状態または取得可能な本文経路を確認できませんでした。
No evidence-backed public reproduction asset is currently recorded.