Affective computing research in Virtual Reality (VR) seeks to develop intelligent systems to identify, understand, and react to human emotions in immersive and interactive Virtual Environment (VE). Various researchers have used facial expressions to c...
Affective computing research in Virtual Reality (VR) seeks to develop intelligent systems to identify, understand, and react to human emotions in immersive and interactive Virtual Environment (VE). Various researchers have used facial expressions to classify emotions in VR. However, several challenges are associated with emotion recognition from the face, such as the VR device hiding half of the face, and a secondary camera is needed to capture the lower face. Therefore, in this study, we developed an innovative system that consists of two parts: a Virtual Plant Shop (VPS) application and a Dual-Attention Emotion Recognition Network (DER-Net). The proposed DER-Net framework leverages audio signals for human emotion recognition and consists of two main phases. In the first phase, the raw audio signals are converted into a visual representation via Mel-Spectrogram (MS). These MSs are then enhanced through filtering and normalization to ensure high-quality input data for training. The second phase used these refined features for DER-Net training and validation. The DER-Net incorporates an optimized EfficientNetV2B2 architecture as a backbone for feature extraction, followed by Spatial Attention (SA) and Channel Attention (CA) modules to further emphasize critical discriminative features while suppressing irrelevant noise. Additionally, we evaluated the DER-Net against various state-of-the-art (SOTA) methods over benchmarks, where our proposed network consistently outperformed SOTA. Moreover, we utilized the model within VPS to recognize the subject’s emotions in real-time and to regulate their emotional state from negativity toward positivity through colors.