← Papers

Paper record

DRSA: Depth-Routed Selective Attention for chili pepper organ segmentation with selective use of estimated monocular depth

Wenhao Zhou · Zixuan Wang · Jianan Chi · Haotian Chen · Pingping Yan · Tiecheng Bai

Plant Methods · 15 Sept 2026 · 10.1186/s13007-026-01581-y

Abstract

Site-specific spraying in chili pepper production requires organ-level segmentation of leaves, peppers, and flowers from handheld field images. However, RGB appearance becomes unreliable under organ overlap, occlusion, dust, and fruit specularity. Offline monocular depth from Depth Anything V2 provides an accessible structural prior without RGB-D sensing, but its reliability is spatially concentrated rather than uniform. To address this, we propose Depth-Routed Selective Attention (DRSA), an estimated-depth-guided segmentation network. DRSA predicts a single per-pixel routing field that, through one shared decision, jointly governs where depth-boundary cross-attention and residual depth fusion contribute, so geometric cues act near organ contours while RGB remains the default carrier. The routing field is calibrated online from depth-on and depth-suppressed predictions without trust-map annotation. We construct PepperField-EstDepth, a self-built dataset of 3, 940 handheld field images paired with estimated monocular depth, on which DRSA achieves \(90.20\%\) mIoU and \(84.48\%\) boundary mIoU, outperforming both RGB-only baselines and attention-based RGB-D fusion baselines; over the RGB segmentation reference, the gains are \(+1.98\) and \(+2.67\) percentage points, respectively. Under group cross-validation, DRSA reaches \(0.8919\pm 0.0031\) mIoU and \(0.8294\pm 0.0046\) boundary mIoU. For the spraying application, DRSA attains a target recall of 0.9814, a target precision of 0.9756, and an organ-level off-target activation of \(2.44\%\) . Visual-pipeline timing shows a segmentation-only latency of 43.0 ms under pre-generated estimated depth, rising to 219.6 ms for the RGB-to-mask visual pipeline when online Depth Anything V2-L depth generation is included. These results support DRSA as a pre-spray organ-level perception module.

Code and data availability

The article describes a self-built dataset (PepperField-EstDepth) and a proposed model (DRSA), but the supplied blocks contain no public deposit, availability statement, or authors' URL for the dataset, code, or trained model. Materials availability states 'Not applicable' and only mentions a third-party pretrained DAv

No evidence-backed public reproduction asset is currently recorded.