← Papers

Paper record

BrinjalFruitX: A field-collected image dataset for machine learning and deep learning-based disease identification in brinjal fruits.

Bitto AK, Hasan MZ, Bijoy MHI, Biplob KBB, Hassan MM, Rana MS, Masum AKM.

Data in brief · 21 Jan 2026 · 10.1016/j.dib.2026.112490

Abstract

Brinjal (Solanum melongena) or eggplant is one of the four most essential vegetable crops that are grown in Bangladesh and contribute significantly to the agricultural industry of the country. Brinjal supports the livelihood of numerous small farmers; however, brinjal is severely susceptible to various fruit diseases, which have serious impacts on yield quality and may cause considerable economic losses. While most existing plant disease datasets primarily focus on leaf-related disorders, only a limited number include fruit-related diseases and even those contain very few classes. This gap is significant because fruit diseases directly affect crop quality, market value, and overall yield. This is why we present here a new and comprehensive dataset that is unparalleled, exclusively for brinjal fruit diseases. This data set consists of 1823 high-quality, labelled images, across five distinct classes: Phomopsis Blight, Shoot and Fruit Borer, Fruit Cracking, Wet Rot, and Healthy Fruit. The images were collected from real farm conditions in numerous areas of Bangladesh to ensure a robust sample of varied environmental and farming practices impacting the growth of diseases. This dataset is designed with the unique aim to support plant disease research and enhance training of deep learning models for autonomous disease detection. Lastly, the dataset will allow early disease detection, enhancing crop management practice, reduction of losses, and increasing farmers' economic returns. The release of this dataset will encourage agricultural research as well as practical use in precision agriculture.

Code and data availability

The paper's brinjal fruit disease image dataset (1823 labeled images, five classes) is publicly deposited on Mendeley Data, and the authors' model training/augmentation code is publicly available on GitHub.

Codepublic

The complete code, along with augmentation scripts and model development, is publicly available in our GitHub repository [12].

Open resource ↗GitHub · html-lines:299-357