With the advancement of Human-Computer Interaction (HCI) and the growing attention on affective computing, Facial Expression Recognition (FER) has been widely utilized in various fields such as healthcare, education, and autonomous driving. However, e...
With the advancement of Human-Computer Interaction (HCI) and the growing attention on affective computing, Facial Expression Recognition (FER) has been widely utilized in various fields such as healthcare, education, and autonomous driving. However, existing studies tend to focus heavily on improving model architectures to enhance accuracy, thereby overlooking the fundamental premise of data quality. In particular, crowdsourced datasets entail label noise problems due to the inherent ambiguity of emotions, and existing methods that reassign labels as single hard labels fail to fundamentally resolve this issue. To address this, this paper proposes a label correction method that utilizes the topological distribution of facial feature vectors to explore the intrinsic structure of data and explicitly accommodates ambiguity by assigning probabilities for each emotion.
The proposed method overcomes low-resolution limitations through image upsampling and then extracts AU and HOG features using a facial feature extraction model.
Subsequently, to resolve dimensional imbalance, individual manifolds are generated using semi-supervised learning-based UMAP and then fused into a single manifold through a topological combination method. In the final stage, the fused manifold is soft-clustered using HDBSCAN, and membership probabilities are calculated based on distance to correct the original hard labels into seven-dimensional soft labels.
As a result of the comparative evaluation, the proposed method improved the Macro F1-Score by 4.40%p compared to the original. Notably, the Macro F1-Scores for 'Disgust' and 'Fear' emotions, which were difficult to classify due to the ambiguity of expression, were significantly improved by 3.70%p and 5.87%p, respectively. Furthermore, the proposed method secured well-balanced performance compared to single-feature models, achieving the lowest variance (0.0144) and standard deviation (0.120) in performance across emotions, thereby demonstrating that it forms uniform and stable predictive decision boundaries for all emotions.