← Papers

Paper record

Monocular Depth Estimation of Tomato in Greenhouse Environment Using the Tomato-MDE Model

Chao Qi · Yawen Cheng · Qian Wu

14 Mar 2025 · 10.36227/techrxiv.174195652.28668799/v1

Abstract

Currently, greenhouse tomato picking robots highly rely on stereo vision cameras to obtain the depth information of the objects. However, due to the inherent flaws of the stereo vision camera, the accuracy of the obtained depth value is extremely susceptible to environmental variations, thus affecting the stability of the picking. In this context, we propose a lightweight monocular depth estimation model for tomatoes (Tomato-MDE) that combines real and virtual datasets, capable of adapting to greenhouse environments with transparent objects. In the encoder, it comprises a simple stacking of the proposed fast self-attention module (FSAM) and depthwise convolution (DWConv), which effectively improves the inference speed of the model. In the decoder, the proposed Integrate module extracts the output of the encoder into image feature representations at different resolutions, and the proposed Merge module then gradually fuses these representations into the final dense depth prediction. The resulting model was tested on 3500 images in the collected greenhouse tomato test dataset, showing that the proposed Tomato-MDE achieved the Depth Error Metric (δ1、δ2、δ3), Absolute Relative Error (Abs Rel) and Root Mean Squared Error (RMSE) of 0.883, 0.934, 0.949, 0.107 and 0.371, with an inference speed of 31.62 FPS under the NVIDIA Tesla V100 GPU environment. In the transparent dataset, Tomato-MDE reached the δ1, δ2, δ3, Abs Rel and RMSE of 0.818, 0.887, 0.907, 0.177 and 0.45, respectively. In addition, this method could be further developed into a perception system for a greenhouse tomato picking robot in the future.

Code and data availability

公開論文であることは確認できましたが、現在の公式API・許可済み取得経路では本文を自動取得できませんでした。

No evidence-backed public reproduction asset is currently recorded.