Paper record
mIT-CMCA: A Cross-Modal Category Alignment Framework for Robust Maize Disease Identification
26 Sept 2025 · 10.21203/rs.3.rs-7286789/v1
Abstract
Abstract Accurate identification of maize diseases is crucial for safeguarding global food security. Traditional image-based methods often struggle with lighting variations, occlusions, and noise, which limits their robustness and generalisation ability. Multimodal approaches that integrate visual and textual information have shown promise. However, these methods frequently require manually curated textual descriptions for each image, increasing data collection costs and limiting scalability and practical implementation. To address these limitations, we proposed a maize image-text framework with Cross-Modal Category Alignment (mIT-CMCA). This approach enforces category-level alignment between image and text modalities, enabling more accurate and interpretable cross-modal mapping. First, we construct cross-modal representations by aligning image and text modalities at the category level within a shared embedding space. Second, inspired by contrastive learning, we introduce a Cross-Modal Category Alignment (CMCA) loss based on category-level textual descriptions, reducing annotation complexity. Finally, we present an Efficient Channel-Spatial Hybrid Attention (CSHA) module that preserves inter-class boundaries with minimal computational overhead to enhance feature discriminability under complex conditions. Experimental results show that the mIT-CMCA model achieves accuracy, precision, recall, and F1-scores of 99.48%, 99.28%, 99.54%, and 99.41%, respectively, on the maize subset of the PlantVillage dataset (MPVD), representing improvements of 0.24%, 0.13%, 0.17%, and 0.15% over the strongest baseline. On the self-built Maize Leaf-Field dataset (MLFD), the model attains corresponding scores of 93.67%, 93.76%, 93.67%, and 93.66%, with respective improvements of 5.34%, 5.13%, 5.34%, and 5.38%.
Code and data availability
The paper uses the public PlantVillage dataset (cited prior work, not paper-specific) and a self-built Maize Leaf-Field dataset (MLFD) plus author code/models, but no block contains any availability statement, deposit, or public URL for the MLFD images, textual descriptions, code, or trained checkpoints. No qualifying,
No evidence-backed public reproduction asset is currently recorded.