← Papers

Paper record

An AI-Driven Dual-Spectral Vision-Language Sensing Framework for Intelligent Agricultural Phenotyping.

Lei Shi · Zhiyuan Chen · Chengze Li · Yang Hu · Xintong Wang · Haibo Wang · Yihong Song

Sensors · 25 Mar 2026 · 10.3390/s26072045

Abstract

Seed varietal purity and physiological viability are critical determinants of crop yield and quality. However, non-destructive assessment faces significant challenges in fine-grained variety discrimination and the perception of internal defects. This study proposes S3-Net, an AI-driven multimodal sensing framework that integrates vision–language alignment with dual-spectral sensor fusion for autonomous seed quality evaluation. We introduce a Knowledge–Vision Alignment (KVA) module that incorporates encyclopedic morphological descriptions to guide feature learning, significantly enhancing few-shot generalization. Complementarily, a Dual-Spectral Fusion (DSF) module combines high-resolution RGB textures with penetrative Short-Wave Infrared (SWIR) sensing to jointly characterize external and internal traits. Experimental results on a custom multimodal dataset of 6000 samples across 12 crop categories demonstrate that S3-Net achieves 96.9% accuracy for species identification and 95.8% for viability detection. Notably, S3-Net outperforms ResNet-50 by 40.3% in extreme 1-shot scenarios. With a stable inference throughput of 95 fps, the system meets the high-throughput demands of industrial-scale applications, providing a robust and efficient solution for intelligent agricultural phenotyping.

Code and data availability

The supplied blocks describe a custom multimodal seed dataset (6005 samples, RGB+SWIR images, text corpus) and the S3-Net model, but contain no public deposit, availability statement, or authors' URL for the dataset, images, code, or trained model. No qualifying paper-specific public assets are present, and allowed URL

No evidence-backed public reproduction asset is currently recorded.