The demand for accurate seabed information has steadily increased due to the growing importance of coastal and marine management and marine spatial planning (MSP). In particular, seabed sediment classification is required for effective marine ecosyste...
The demand for accurate seabed information has steadily increased due to the growing importance of coastal and marine management and marine spatial planning (MSP). In particular, seabed sediment classification is required for effective marine ecosystem conservation, seabed habitat mapping, and coastal development. However, grab samples, which provide ground-truth information for seabed sediment classification are difficult to acquire over wide marine areas and often produce an imbalanced class distribution, which reduces the robustness of the classifications derived from machine learning models. The present study thus combined grab-sample augmentation based on simple linear iterative clustering with confidence-based auto-labeling (AL) to mitigate the training constraints arising from the limited number of grab samples available for seabed sediment classification. A total of 23 variables were extracted from multibeam echosounder data acquired from the East Sea of Korea, including bathymetry, backscatter, and gray-level co-occurrence matrix-based texture features. The classification performance of four machine-learning models (random forest, extra trees, light gradient boosting machine, and extreme gradient boosting) was evaluated based on grab samples under four Folk-based classification schemes (Schemes 1–4). Scheme 1 achieved an overall accuracy (OA) of approximately 88.82%, while Scheme 4 achieved an OA of approximately 84.11%. Notably, when AL was applied, all four schemes exhibited an average performance improvement of about 10–17 percentage points. In addition, the gravel class, which had a user accuracy of 0% in the original dataset, improved to 66.67% when combining data augmentation at an augmentation ratio of 5 with AL, indicating that the class imbalance was alleviated. These findings demonstrate that high-accuracy seabed sediment classification is achievable with limited in-situ samples by integrating data augmentation with AL. The proposed approach is thus expected to produce important baseline data for seabed habitat mapping, MSP, marine ecosystem conservation, and marine resource management.