Recent advances in artificial intelligence (AI)–based visual recognition have rapidly expanded into military, medical, and ecological domains, becoming a core technology for real-time surveillance, reconnaissance, and automatic target recognition in...
Recent advances in artificial intelligence (AI)–based visual recognition have rapidly expanded into military, medical, and ecological domains, becoming a core technology for real-time surveillance, reconnaissance, and automatic target recognition in modern warfare. Among these applications, Camouflaged Object Detection (COD) aims to identify objects that visually blend into their surroundings, a task that remains one of the most challenging yet valuable problems in defense. Although recent COD models incorporating multi-scale feature extraction, attention mechanisms, and transformer-based architectures have demonstrated significant performance improvements, detecting camouflaged objects in realistic and complex environments remains unresolved. The emergence of large-scale public datasets has improved training efficiency and facilitated research that simulates realistic operational scenarios. However, despite these advances, COD in real-world environments remains unresolved.
In operational settings, salien0t objects (SO) such as vehicles, civilians, or signboards often coexist with camouflaged soldiers or equipment, introducing significant confusion for detection models. Since conventional COD networks focus solely on “detecting camouflaged objects,” they often fail to distinguish non-target salient objects, leading to false detections that degrade operational reliability. For example, a bright civilian vehicle in surveillance footage may be falsely identified as a camouflaged target. Moreover, existing COD datasets primarily emphasize camouflaged objects, lacking complex coexistent scenes that reflect real battlefields, which ultimately limits the model’s generalization and robustness across environments.
To overcome these limitations, this study proposes a learning framework that integrates synthetic data generation with negative learning. Using CamDiff, a latent diffusion–based algorithm that masks regions of a background image, restores them through the Latent Diffusion Model (LDM), and validates realism via Contrastive Language–Image Pretraining (CLIP), we synthesized realistic scenes containing both camouflaged and salient objects. These synthetic images were then used to construct a synthetic dataset, providing a diverse training environment. During training, salient objects were explicitly labeled as non-detection targets, enabling the model to learn “what not to detect” through negative learning. The baseline model, PFNet (Positioning and Focusing Network), was used for evaluation under varying data configurations.
Experimental results demonstrate that models trained on synthetic data achieved a 16.6% reduction in MAE, with corresponding improvements across S-, E-, and F-measures. Furthermore, incorporating salient object images through negative learning led to an additional 4.6% reduction in MAE and performance gains of +2.6% in E-measure and +3.2% in F-measure, confirming a distinct false-positive suppression effect. In summary, the proposed approach not only enhances detection accuracy but also strengthens reliability in complex military environments by explicitly teaching the model to disregard non-target objects. This framework offers a practical step toward more trustworthy and robust AI-based surveillance and reconnaissance systems. By applying CamDiff-based synthetic data to military scenarios, this study also alleviates the limitations of collecting real-world defense data. The proposed approach is expected to contribute to future defense-oriented COD analysis and enhance the reliability of AI-based battlefield perception systems.