Paper record
A dynamic multi-scale feature fusion and hierarchical attention network for leaf segmentation
Discover Computing · 3 Sept 2026 · 10.1007/s10791-026-10494-2
Abstract
Abstract Accurate leaf instance segmentation is fundamental to quantifying morphological and structural traits in image-based plant phenotyping. However, substantial variations in leaf scale, shape, and orientation, together with dense overlap and occlusion, often lead to ambiguous instance boundaries. In addition, repeated downsampling and feature reconstruction can progressively erode fine structural details, hindering contour preservation and the separation of adjacent leaves. To address these interrelated challenges, we propose DMSHA-Net, a dynamic multi-scale feature fusion and hierarchical attention network for leaf instance segmentation. Its direction-aware Multi-Scale Feature Aggregation (MSFA) encoder captures complementary horizontal and vertical contextual information across multiple receptive-field scales, thereby improving the representation of diverse leaf morphologies. The Dense Feature Aggregation (DFA) decoder selectively integrates deep semantic and shallow structural features through stage-specific attention mechanisms, where self-attention at coarse resolutions models long-range dependencies, whereas lightweight channel recalibration at high resolutions refines local structures and boundaries. The Learnable Feature Fusion (LFF) module subsequently integrates multi-level semantic and boundary features using normalized learned weights. DMSHA-Net achieves Best Dice (BD) scores of 93.17%, 85.12%, and 90.91% on KOMATSUNA, MSU-PID, and the CVPPP-A1 subset, respectively, with corresponding foreground–background Dice (FBD) scores of 98.19%, 91.02%, and 98.24%. These results indicate that DMSHA-Net achieves competitive segmentation performance across the three datasets, particularly in scenarios involving pronounced scale variation and dense leaf overlap.
Code and data availability
The paper's own source code is not publicly deposited; the authors state it is available only upon reasonable request. The evaluation datasets (CVPPP, KOMATSUNA, MSU-PID) are public benchmarks from cited prior work rather than paper-specific assets, and no trained models or author analysis code with a public URL are提供的
No evidence-backed public reproduction asset is currently recorded.