Object detection and segmentation in complex environments remain challenging tasks, particularly when dealing with multimodal data and small object detection. This dissertation presents a series of novel approaches to improve detection accuracy across...
Object detection and segmentation in complex environments remain challenging tasks, particularly when dealing with multimodal data and small object detection. This dissertation presents a series of novel approaches to improve detection accuracy across different domains, including maritime surveillance, handwritten character recognition, nanoparticle detection, and pedestrian detection in low-light conditions. The research primarily focuses on multimodal fusion and anchorless detection frameworks, offering advanced solutions tailored to each scenario.
First, an efficient multimodal fusion method is introduced for ship detection and tracking in electro-optical (EO) and infrared (IR) imagery. By leveraging paired sequence frames and refining the fusion process, the model enhances detection robustness in challenging maritime conditions. Next, Gaussian heatmap-based object localization is explored in the domain of handwritten character recognition, demonstrating how anchor-free approaches can improve detection precision for complex scripts.
The study further extends to nanoparticle detection in microscopy images, addressing small object detection challenges through advanced feature extraction techniques. Finally, a refined fusion and anchorless detection framework is applied to pedestrian detection in near-infrared (NIR) and depth imagery, improving detection accuracy in low-light and occluded environments.
Throughout this work, we introduce novel augmentation techniques, multimodal feature fusion strategies, and anchorless detection models that significantly enhance object detection performance. Experimental results across multiple datasets confirm the effectiveness of the proposed approaches compared to state-of-the-art methods. The findings contribute to the broader field of computer vision by advancing detection methodologies in diverse and challenging imaging scenarios.