This paper proposes a temporal multi-view camera-based 3D object detection framework robust to ego-pose estimation uncertainty by utilizing vehicle kinematics information derived from the in-vehicle Controller Area Network (CAN). Temporal 3D object de...
This paper proposes a temporal multi-view camera-based 3D object detection framework robust to ego-pose estimation uncertainty by utilizing vehicle kinematics information derived from the in-vehicle Controller Area Network (CAN). Temporal 3D object detection models typically fuse features from past frames with the current view to compensate for occlusion and sparse observations. This process requires inter-frame alignment, which generally relies on the ego-pose estimated from external sensors such as IMU and GPS. However, inaccurate pose information leads to the misalignment of historical features, resulting in accumulated errors and significant degradation in detection performance.
To mitigate alignment errors caused by such pose uncertainty, this study leverages additional vehicle kinematics information. Specifically, we estimate the relative translation and rotation changes between frames using CAN-based kinematics, which are accessible without additional high-precision sensors. These estimates are then integrated as residuals into the ego-transform stage of the temporal multi-view 3D object detection model. The proposed method ensures that information from past frames is stably aligned with the current coordinate system even when external ego-pose data is corrupted or unreliable, thereby preventing catastrophic degradation in temporal detection performance.
To validate the effectiveness of the proposed method, we conducted comparative experiments against the baseline(StreamPETR) using the nuScenes dataset under evaluation scenarios with artificially injected ego-pose noise. Experimental results demonstrate that the proposed model achieves a relative performance improvement of 11.5% in mAP and 8.7% in NDS compared to the baseline under noisy pose conditions, quantitatively proving its robustness against pose uncertainty. In conclusion, this study shows that the stability of temporal multi-view camera-based 3D object detection can be enhanced solely using internal vehicle kinematics, mitigating the reliance on expensive high-precision external sensors.