This study is aimed at proposing an integrated network to perform two tasks: object detection and depth estimation in autonomous driving and robotics. The approach combines a bird’s eye view-based candidate generation module with an image-point clou...
This study is aimed at proposing an integrated network to perform two tasks: object detection and depth estimation in autonomous driving and robotics. The approach combines a bird’s eye view-based candidate generation module with an image-point cloud, cross-attention fusion structure to exploit complementary spatial and visual cues from both modalities. Moreover, an input-dependent query initialization module is employed to initiate detection in likely object regions, thereby reducing unnecessary candidates. To improve depth accuracy, Hungarian matching is applied, and performance is quantitatively evaluated by using the root mean square error. Experiments on the KITTI dataset demonstrated that the method achieved superior performance over existing approaches involving cars, pedestrians, and cyclists. These results indicate that the proposed network can provide robust and precise perception even in complex driving environments.