SAMFNet: Scene-aware sampling and multi-stage fusion for multimodal 3D object detection

Name: SAMFNet: Scene-aware sampling and multi-stage fusion for multimodal 3D object detection
Keywords: Scene-aware sampling, Virtual point clouds, Multimodal 3D object detection, TA1-2040, Engineering (General). Civil engineering (General), Multi-stage feature fusion

descriptionPublicationkeyboard_double_arrow_right Article 01 Jul 2025 English Publisher:Elsevier BVJournal:Alexandria Engineering Journal, volume 126, pages 90-104 (issn: 1110-0168,

Authors: Baotong Wang; Chenxing Xia; Xiuju Gao; Bin Ge; Kuan-Ching Li; Xianjin Fang; Yan Zhang; +1 Authors

doi: 10.1016/j.aej.2025.03.129

SAMFNet: Scene-aware sampling and multi-stage fusion for multimodal 3D object detection

- Summary
- Subjects
- Metrics

Abstract

Recently, multimodal 3D object detection (M3OD) that fuses the complementary information from LiDAR data and RGB images has gained significant attention. However, the inherent structural differences between point clouds and images pose fusion challenges, significantly hindering the exploration of correlations within multimodal data. To address this issue, this paper introduces an enhanced multimodal 3D object detection framework (SAMFNet), which leverages virtual point clouds generated from depth completion. Specifically, we design a scene-aware sampling module (SASM) that employs tailored sampling strategies for different bins based on the density distribution of point clouds. This effectively alleviates the detection bias problem while ensuring the key information of virtual points, significantly reducing the computational cost. In addition, we introduce a multi-stage feature fusion module (MSFFM) that embeds point-level and regional-adaptive feature fusion strategies to generate more informative multimodal features by fusing features with different granularities. To further improve the accuracy of model detection, we also introduce a confidence prediction branch unit (CPBU), which improves the detection accuracy by predicting the confidence of feature classification in the intermediate stage. Extensive experiments on the challenging KITTI dataset demonstrate the validity of our model.

Keywords

Scene-aware sampling, Virtual point clouds, Multimodal 3D object detection, TA1-2040, Engineering (General). Civil engineering (General), Multi-stage feature fusion

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

Average

gold