Global Fusion and Progressive Decoding Network for RGB-D Salient Object Detection
WEI Longsheng1,2,3, MI Yichen1,3, ZHOU Zhanxiang1,3
1. School of Artificial Intelligence and Automation, China University of Geosciences, Wuhan 430074; 2. Yunnan Province Key Laboratory of Low Light Detection and Intelligent Visual Navigation, Yunnan Kunming 650214; 3. Hubei Key Laboratory of Advanced Control and Intelligent Automation for Complex Systems, China University of Geosciences, Wuhan 430074
Abstract:Existing RGB-D salient object detection methods predominantly employ homogenized fusion strategies and neglect the inherent discrepancies across spatial and semantic dimensions. To address these issues, a global fusion and progressive decoding network for RGB-D salient object detection is proposed in this paper. An asymmetric dual-stream encoder architecture is utilized, and differentiated backbone networks tailored to the distinct information densities of RGB and depth modalities are designed. First, a detail-aware module is developed to replace traditional symmetric fusion. Global RGB texture features are employed as spatial guidance for the asymmetric filtering of low-quality depth features with fine boundaries preserved and sensor noise suppressed. Then, a semantic aggregation module is developed to construct parallel multi-scale receptive fields. A channel attention mechanism is introduced for cross-modal semantic purification, optimizating precise localization and semantic alignment for multi-scale targets. Finally, a refined decoding module is constructed to cooperate with large receptive fields to extract channel attention through a dual-branch residual structure, effectively compensating for information loss and achieving a progressive reconstruction of salient objects from high-level semantics to low-level details. Experiments on six public datasets demonstrate that the proposed network achieves superior detection performance.
[1] 化春键,姚烨涛,蒋毅,等.特征增强和渐进式解码的RGB-D显著性检测[J].计算机科学与探索, 2025, 19(9): 2419-2429. (Hua C J, Yao Y T, Jiang Y, et al. RGB-D salient object detection with feature enhancement and progressive decoding[J]. Journal of Frontiers of Computer Science and Technology, 2025, 19(9): 2419-2429.) [2] Itti L, Koch C, Niebur E. A model of saliency-based visual attention for rapid scene analysis[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 1998, 20(11): 1254-1259. [3] Wei L S, Zong G Y. EGA-Net: Edge feature enhancement and glo-bal information attention network for RGB-D salient object detection[J]. Information Sciences, 2023, 626: 223-248. [4] Zhou W J, Guo Q L, Lei J S, et al. IRFR-Net: Interactive recursive feature-reshaping network for detecting salient objects in RGB-D images[J]. IEEE Transactions on Neural Networks and Learning Systems, 2025, 36(3): 4132-4144. [5] Wu Z W, Paudel D P, Fan D P, et al. Source-free depth for object pop-out[C]//Proceedings of the IEEE/CVF International Confe-rence on Computer Vision. Washington, USA: IEEE, 2023: 1032-1042. [6] Feng G, Meng J Y, Zhang L H, et al. Encoder deep interleaved network with multi-scale aggregation for RGB-D salient object detection[J/OL]. Pattern Recognition, 2022, 128. https://doi.org/10.1016/j.patcog.2022.108666. [7] Fang X, Jiang M F, Zhu J C, et al. M2RNet: multi-modal and multi-scale refined network for RGB-D salient object detection[J/OL]. Pattern Recognition, 2023, 135. https://doi.org/10.1016/j.patcog.2022.109139. [8] Jin D Z, Shao F, Xie Z X, et al. CAFCNet: cross-modality asy-mmetric feature complement network for RGB-T salient object detection[J/OL]. Expert Systems with Applications, 2024, 247. https://doi.org/10.1016/j.eswa.2024.123222. [9] Wang F S, Wang R M, Sun F M. DCMNet: discriminant and cross-modality network for RGB-D salient object detection[J/OL]. Expert Systems with Applications, 2023, 214. https://doi.org/10.1016/j.eswa.2022.119047. [10] Wu J Y, Sun F M, Xu R, et al. Aggregate interactive learning for RGB-D salient object detection[J/OL]. Expert Systems with App-lications, 2022, 195. https://doi.org/10.1016/j.eswa.2022.116614. [11] Jiang M F, Ma J H, Chen J T, et al. PATNet: patch-to-pixel attention-aware transformer network for RGB-D and RGB-T salient object detection[J/OL]. Knowledge-Based Systems, 2024, 291. https://doi.org/10.1016/j.knosys.2024.111597. [12] Wu Y H, Liu Y, Xu J, et al. MobileSal: extremely efficient RGB-D salient object detection[J]. IEEE Transactions on Pattern Ana-lysis and Machine Intelligence, 2022, 44(12): 10261-10269. [13] Zeng Z H, Liu H J, Chen F L, et al. AirSOD: a lightweight network for RGB-D salient object detection[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2024, 34(3): 1656-1669. [14] Zhou W J, Zhu Y, Lei J S, et al. LSNet: lightweight spatial boosting network for detecting salient objects in RGB-thermal images[J]. IEEE Transactions on Image Processing, 2023, 32: 1329-1340. [15] Jin X, Yi K, Xu J. MoADNet: mobile asymmetric dual-stream networks for real-time and lightweight RGB-D salient object detection[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2022, 32(11): 7632-7645. [16] Zhang W B, Ji G P, Wang Z, et al. Depth quality-inspired feature manipulation for efficient RGB-D salient object detection[C]//Proceedings of the 29th ACM International Conference on Multimedia. New York, USA: ACM, 2021: 731-740. [17] Chen H, Li Y F. Progressively complementarity-aware fusion network for RGB-D salient object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2018: 3051-3060. [18] Chen Z Y, Cong R M, Xu Q Q, et al. DPANet: depth potentiality-aware gated attention network for RGB-D salient object detection[J]. IEEE Transactions on Image Processing, 2021, 30: 7012-7024. [19] Huang N C, Yang Y, Zhang D W, et al. Employing bilinear fusion and saliency prior information for RGB-D salient object detection[J]. IEEE Transactions on Multimedia, 2022, 24: 1651-1664. [20] Wan B, Zhou X F, Sun Y Q, et al. MFFNet: multi-modal feature fusion network for V-D-T salient object detection[J]. IEEE Tran-sactions on Multimedia, 2024, 26: 2069-2081. [21] Pang Y W, Zhao X Q, Zhang L H, et al. CAVER: cross-modal view-mixed transformer for bi-modal salient object detection[J]. IEEE Transactions on Image Processing, 2023, 32: 892-904. [22] Wang R M, Wang F S, Su Y M, et al. Attention-guided multi-modality interaction network for RGB-D salient object detection[J]. ACM Transactions on Multimedia Computing, Communications and Applications, 2023, 20(3): 1-22. [23] Wang W H, Xie E Z, Li X, et al. Pyramid vision transformer: a versatile backbone for dense prediction without convolutions[C]//Proceedings of the IEEE/CVF International Conference on Compu-ter Vision. Washington, USA: IEEE, 2021: 548-558. [24] Zheng Z H, Wang P, Liu W, et al. Distance-IoU loss: faster and better learning for bounding box regression[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2020, 34(7): 12993-13000. [25] Fan D P, Gong C, Cao Y, et al. Enhanced-alignment measure for binary foreground map evaluation[C]//Proceedings of the 27th International Joint Conferences on Artificial Intelligence. San Francisco, USA: IJCAI, 2018: 698-704. [26] Fan D P, Cheng M M, Liu Y, et al. Structure-measure: a new way to evaluate foreground maps[C]//Proceedings of the IEEE International Conference on Computer Vision. Washington, USA: IEEE, 2017: 4558-4567. [27] Achanta R, Hemami S, Estrada F, et al. Frequency-tuned salient region detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2009: 1597-1604. [28] Perazzi F, Krähenbühl P, Pritch Y, et al. Saliency filters: contrast based filtering for salient region detection[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Washington, USA: IEEE, 2012: 733-740. [29] Jin W D, Xu J, Han Q, et al. CDNet: complementary depth network for RGB-D salient object detection[J]. IEEE Transactions on Image Processing, 2021, 30: 3376-3390. [30] Chen Q, Liu Z, Zhang Y, et al. RGB-D salient object detection via 3D convolutional neural networks[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2021, 35(2): 1063-1071. [31] Fan D P, Zhai Y J, Borji A, et al. BBS-Net: RGB-D salient object detection with a bifurcated backbone strategy network[C]//Proceedings of the 16th European Conference on Computer Vision. Berlin, Germany: Springer, 2020: 275-292. [32] Zhao X Q, Pang Y W, Zhang L H, et al. Self-supervised pretrai-ning for RGB-D salient object detection[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2022, 36(3): 3463-3471. [33] Cong R M, Lin Q W, Zhang C, et al. CIR-Net: cross-modality interaction and refinement for RGB-D salient object detection[J]. IEEE Transactions on Image Processing, 2022, 31: 6800-6815. [34] Bi H B, Wu R W, Liu Z Q, et al. Cross-modal hierarchical interaction network for RGB-D salient object detection[J/OL]. Pattern Recognition, 2023, 136. https://doi.org/10.1016/j.patcog.2022.109194. [35] Cong R M, Liu H Y, Zhang C, et al. Point-aware interaction and CNN-induced refinement network for RGB-D salient object detection[C]//Proceedings of the 31st ACM International Conference on Multimedia. New York, USA: ACM 2023: 406-416. [36] Wu J S, Hao F W, Liang W Y, et al. Transformer fusion and pi-xel-level contrastive learning for RGB-D salient object detection[J]. IEEE Transactions on Multimedia, 2024, 26: 1011-1026. [37] Song P P, Li W Y, Zhong P Y, et al. Synergizing triple attention with depth quality for RGB-D salient object detection[J/OL]. Neurocomputing, 2024, 589. https://doi.org/10.1016/j.neucom.2024.127672. [38] Hu X H, Sun F M, Sun J, et al. Cross-modal fusion and progre-ssive decoding network for RGB-D salient object detection[J]. International Journal of Computer Vision, 2024, 132(8): 3067-3085. [39] Duan S S, Yang X, Wang N N, et al. Lightweight RGB-D salient object detection from a speed-accuracy tradeoff perspective[J]. IEEE Transactions on Image Processing, 2025, 34: 2529-2543. [40] Han J Y, Wang M Y, Wu W Y, et al. Perceptual localization and focus refinement network for RGB-D salient object detection[J/OL]. Expert Systems with Applications, 2025, 259. https://doi.org/10.1016/j.eswa.2024.125278. [41] Sun F M, Ren P, Yin B W, et al. CATNet: a cascaded and aggregated transformer network for RGB-D salient object detection[J]. IEEE Transactions on Multimedia, 2024, 26: 2249-2262.