Industrial visual inspection is often constrained by an acute scarcity of labeled defective examples and by large variability in defect appearance across different products and imaging conditions, which together undermine the generalization of supervi...
Industrial visual inspection is often constrained by an acute scarcity of labeled defective examples and by large variability in defect appearance across different products and imaging conditions, which together undermine the generalization of supervised detectors. Our proposed system addresses these challenges through an integrated, data-and-model approach: a diffusion-based generative augmentor is used to synthesize realistic defective instances that expand and diversify the anomaly class, while a multi-branch convolutional backbone whose outputs are adaptively fused via a transformer-based attention module provides complementary local and global feature representations. The combined effect of targeted synthetic augmentation and attention-guided feature fusion reduces overfitting to limited defect samples and increases the system’s sensitivity to both subtle texture changes and larger structural anomalies. We validate the proposed pipeline on established industrial benchmarks and perform comprehensive ablation studies to isolate the contribution of each component. Empirical results demonstrate strong discriminative performance, achieving an AUROC of 98.15% on the MVTec LOCO AD dataset and 97.56% on the MVTec AD dataset, while consistently improving robustness under pronounced class imbalance and varying imaging conditions.