Accurate classification of crop physiological disorders is essential for providing farmers with appropriate agricultural countermeasures and preventing yield loss. Traditionally, physiological measurements and laboratory tissue analyses have been used...
Accurate classification of crop physiological disorders is essential for providing farmers with appropriate agricultural countermeasures and preventing yield loss. Traditionally, physiological measurements and laboratory tissue analyses have been used to detect such disorders. While these methods can accurately identify physiological stresses by measuring crop variables such as stomatal conductance and photosynthetic efficiency, they are labor-intensive and require specialized equipment, making them impractical for large-scale field monitoring and real-time decision-making for farmers. As a result, hyperspectral remote sensing has emerged as a promising alternative for detecting stress signals in wide-area farmland in real time, and has become a focus of active research. However, current research on detecting and classifying crop physiological disorders using hyperspectral remote sensing faces two main challenges. First, most studies have focused on binary classification, distinguishing only between healthy and unhealthy plants. This approach is limited in practical decision-support under real-world complex stress environments, such as those caused by climate change, encountered in actual farms. Second, most current stress-detection models do not address the class imbalance issue that arises when detecting early-stage stress symptoms, which are typically not visually apparent. This leads to inefficient model training.
To address these two challenges, this study conducted multi-class stress classification in lettuce by applying five different stress factors (four abiotic: acidity, drought, salinity, and herbicide; and one biotic: bacterial soft rot caused by Pectobacterium carotovorum). In addition, before visual symptoms appeared, we applied a method that statistically separates stress-affected areas in the spectral data of each treatment group, using the healthy area of a control group as a reference, to determine if this improves model performance. The acquired data included hyperspectral images and physiological indicators measured using a porometer, collected from October 30 to November 7, 2023. Three machine learning models—Random Forest, XGBoost, and Support Vector Machine (SVM)—were trained, and model performance was evaluated using accuracy, precision, recall, and F1-score. For Random Forest and XGBoost, which provide feature importance maps, a combination of optimal wavelength bands with minimal performance loss was selected. Based on these results, a new multispectral classification model was developed that required a reduced number of spectral bands. Physiological indicators measured with the porometer showed statistically significant (p<0.3) differences on the 8th day (November 7) after stress treatment. By applying the stress-area extraction method, classification accuracy improved significantly compared to the conventional approach, increasing from 21% to 64% for Random Forest, 23% to 60% for XGBoost, and 22% to 53% for SVM. Other performance metrics calculated from confusion matrices also improved across the board. Feature importance analysis for the efficient multispectral model revealed that for Random Forest, wavelengths around 500 nm, 620 nm, 680 nm, 700 nm, and 750 nm contributed the most, while for XGBoost, wavelengths around 570 nm, 620 nm, 700 nm, and 820 nm were most significant. Furthermore, the optimized multispectral model achieved comparable classification performance—a reduction in required bands by up to 81% (from 110 bands to 21 bands), with only a 4% decrease in accuracy (from 64% to 60%). This demonstrates the feasibility of developing high-performing multispectral classification models with minimal wavelength requirements.