Current speech recognizers have advanced to the point that they can recognize continuous speech containing many vocabularies with high accuracy. Despite these achievements, however, the performances of speech recognizers drop rapidly in actual uses wh...
Current speech recognizers have advanced to the point that they can recognize continuous speech containing many vocabularies with high accuracy. Despite these achievements, however, the performances of speech recognizers drop rapidly in actual uses where out-of-vocabulary words and environmental distortions like noises and indoor reverberations occur frequently. Many attempts at the robust speech recognition were able to improve the performance under certain environmental circumstances, but they have yet to achieve consistent and reliable recognition performance across various environments.
The major difficulty in the robust speech recognition is that we cannot have any prior knowledge about out-of-vocabulary words or environmental distortions. Finding the characteristics of vocabulary sounds that are robust against distortions is yet another difficult problem. In this thesis, a new robust similarity measure for the sounds with distinctive spectral peaks will be introduced. Examples of such sounds include vowel sounds and mechanical sounds like doorbell sounds. By restricting the target vocabulary sounds as those having spectral peaks and utilizing them, the target sounds that are corrupted by distortions can be recognized, provided that their characteristic peaks have not totally nullified by distortions. And the characteristic spectral peaks that are nearly unique to each sound can be used to reject the out-of-vocabulary sounds which do not possess such peaks.
The performance of proposed similarity measure was tested by two separate experiments. The first experiment is mechanical sound recognition in various home environments. The evaluation of the proposed method showed 99.7% sound recognition accuracy as well as 99.7% noise rejection accuracy. Extension to multiple, concurrent sound recognition achieved 92.8% accuracy for multiple sounds, while retaining the noise rejection accuracy of 98.9% and single sound accuracy of 97.5%.
Another experiment is the voice activity detection (VAD) experiment using vowel sounds. The system was trained using vowel sounds of single female speaker from the TIMIT database, and tested with the utterances of other speakers with varying SNR levels. The performance analysis and comparison with other VAD algorithms will be provided.