RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    검색결과 좁혀 보기

    선택해제
    • 좁혀본 항목 보기순서

      • 원문유무
      • 음성지원유무
      • 원문제공처
        펼치기
      • 등재정보
        펼치기
      • 학술지명
        펼치기
      • 주제분류
        펼치기
      • 발행연도
        펼치기
      • 작성언어

    오늘 본 자료

    • 오늘 본 자료가 없습니다.
    더보기
    • 무료
    • 기관 내 무료
    • 유료
    • Personality of Synthetic Speech: A Literature Review

      Dawoon Jeong,Sung H. Han 대한인간공학회 2018 대한인간공학회 학술대회논문집 Vol.2018 No.5

      The aim of this study is to search for design factors of personality of synthetic speech and its preferred personality through a literature review. Finding the factors affecting the personality of synthetic speech is important because the synthetic speech is now widely used. Two categories of design factors were found through the literature survey: speech parameters and speech synthesis techniques. Speech parameters include pitch, pitch range, frequency range, speaking rate, intensity, fillers, length and frequency of pauses, and voice quality. Speech synthesis techniques have two typical strategies, unit selection synthesis and statistical parametric synthesis. Using these factors, we can design a variety of synthetic speech with personality such as extraversion, neuroticism, agreeableness, conscientiousness, and openness. Preference of synthetic speech depends on personality of synthetic speech and personality of the users. This research can be used to design synthetic speech for making desirable persona and making voice of speech-impaired person.

    • KCI등재

      Survey on Deep Learning-based Speech Technologies in Voice Chatbot Systems

      Seunghee Ma,Junseok Oh,Minseo Kim,김지환 한국인터넷정보학회 2025 KSII Transactions on Internet and Information Syst Vol.19 No.5

      Recent advancements in large language models (LLMs) such as ChatGPT have contributed to the development of chatbot systems. Specifically, speech has been recognized as the optimal tool for interactive dialogue, leading to increased interest in voice chatbots. Voice chatbots offer information and services through voice interactions, enhancing user experience and improving service accessibility. This survey paper introduces the latest developments in the core technologies of voice chatbot systems, including deep learning-based automatic speech recognition, speech synthesis, and speech emotion recognition. It focuses on advanced research to enhance speed and performance, which is crucial for applying speech technologies in voice chatbots. In automatic speech recognition, we introduce methodologies such as Connectionist Temporal Classification (CTC), Attention based Encoder-Decoder (AED), and Recurrent Neural Network Transducer (RNN-T), along with studies optimizing Transformer-based models for real-time automatic speech recognition. In speech emotion recognition, we explore the use of pre-trained models and the latest techniques for accurate emotion prediction. For speech synthesis, the focus extends to two-stage and End-to-End (E2E) approaches, with additional research on integrating emotional information to generate natural speech.

    • KCI등재

      천이구간 추출 및 근사합성에 의한 음성신호 압축과 복원

      이광석,이병로,Lee, Kwang-Seok,Lee, Byeong-Ro 한국정보통신학회 2009 한국정보통신학회논문지 Vol.13 No.2

      In a speech coding system using excitation source of voiced and unvoiced, it would be involved a distortion of speech qualify in case coexist with a voiced and an unvoiced consonants in a frame. So, We proposed TS(Transition Segment) including unvoiced consonant searching and extraction method in order to uncoexistent with a voiced and unvoiced consonants in a frame. This research present a new method of TS approximate-synthesis by using Least Mean Square and frequency band division. As a result, this method obtain a high qualify approximation-synthesis waveforms within TS by using frequency information of 0.547kHz below and 2.813kHz above. The important thing is that the maximum error signal can be made with low distortion approximation-synthesis waveform within TS. This method has the capability of being applied to a new speech coding of Voiced/Silence/TS, speech analysis and speech synthesis. 유 무성음의 음원을 이용한 음성부호화 시스템에서는 프레임 내에 유성자음과 무성자음이 공존하는 경우에 음질왜곡을 일으킬 수 있다. 따라서 프레임 내에 유성자음과 무성자음이 공존하지 않도록 하기 방법으로써 무성자음을 탐색하고 검출을 포함하는 천이 구간을 제안하였다. 본 연구는 최소 자승법과 주파수 대 역 분할법을 사용함으로써 TS 근사합성의 새로운 방식을 제시하였으며 결과적으로 이는 0.547KHz이하와 2.813kHz 이상에서의 주파수 정보를 이용함으로써 TS내에서 고품질의 근사합성 파형을 얻을 수 있었다. 보다 중요한 것은 최대 오류신호는 TS 내에 저 왜곡 근사 합성파형이 생길 수 있다는 것이다. 이 방식은 유성음/묵음/TS의 새로운 음성부호화, 음성해석 및 음성 합성에 적용할 수 있으리라 생각한다.

    • KCI등재

      기본주파수와 성도길이의 상관관계를 이용한 HTS 음성합성기에서의 목소리 변환

      유효근(Yoo, Hyogeun),김영관(Kim, Younggwan),서영주(Suh, Youngjoo),김회린(Kim, Hoirin) 한국음성학회 2017 말소리와 음성과학 Vol.9 No.1

      The main advantage of the statistical parametric speech synthesis is its flexibility in changing voice characteristics. A personalized text-to-speech(TTS) system can be implemented by combining a speech synthesis system and a voice transformation system, and it is widely used in many application areas. It is known that the fundamental frequency and the spectral envelope of speech signal can be independently modified to convert the voice characteristics. Also it is important to maintain naturalness of the transformed speech. In this paper, a speech synthesis system based on Hidden Markov Model(HMM-based speech synthesis, HTS) using the STRAIGHT vocoder is constructed and voice transformation is conducted by modifying the fundamental frequency and spectral envelope. The fundamental frequency is transformed in a scaling method, and the spectral envelope is transformed through frequency warping method to control the speaker’s vocal tract length. In particular, this study proposes a voice transformation method using the correlation between fundamental frequency and vocal tract length. Subjective evaluations were conducted to assess preference and mean opinion scores(MOS) for naturalness of synthetic speech. Experimental results showed that the proposed voice transformation method achieved higher preference than baseline systems while maintaining the naturalness of the speech quality.

    • KCI등재

      중국어 음성합성: 진단과 과제

      이옥주 ( Lee Ok Joo ) 한국중국어문학회 2021 중국문학 Vol.106 No.-

      Recent years have witnessed the remarkable improvement of speech synthesis technology, which is a fundamental component of numerous AI programs. Chinese text-to-speech technology, which has been also rapidly improved, is used in a variety of programs for language teaching and learning as well as automatic translation. This paper examines major problems with Chinese synthesized speech in several widely-used programs, and discusses the prosodic modeling and the types of speech database that can be used to enhance the naturalness of synthesized speech sounds. It further notes that an increasing body of interdisciplinary research between phonetics and engineering may help to develop new analytic methods in the field of Chinese phonology and phonetics.

    • KCI등재

      양방향 상태 공간 모델과 감정 유도 교차 주의를 활용한 감정 강도 제어 음성 합성

      함인성,오경석,송락빈,구본화,고한석 한국음향학회 2025 한국음향학회지 Vol.44 No.5

      Recent advances have led to the development of emotion-intensity controllable speech synthesis models. However, these systems often suffer from degraded speech quality and unnatural emotional expressions, creating a critical gap between human-like expressiveness and synthetic speech. To address these challenges, we propose a novel framework that replaces traditional Transformer architectures with Bidirectional State Space Models for emotion-intensity controllable speech synthesis. Our approach incorporates an Emotion-Guided Cross Attention mechanism to effectively model interactions between emotional and acoustic characteristics, enhancing fine-grained intensity control, speech quality, and naturalness. Experimental results demonstrate that this approach achieves comparable or better performance than existing systems in terms of speech naturalness.

    • KCI등재

      RawNet3를 통해 추출한 화자 특성 기반 원샷 다화자 음성합성 시스템

      한소희,엄지섭,김회린 한국음성학회 2024 말소리와 음성과학 Vol.16 No.1

      Recent advances in text-to-speech (TTS) technology have significantly improved the quality of synthesized speech, reaching a level where it can closely imitate natural human speech. Especially, TTS models offering various voice characteristics and personalized speech, are widely utilized in fields such as artificial intelligence (AI) tutors, advertising, and video dubbing. Accordingly, in this paper, we propose a one-shot multi-speaker TTS system that can ensure acoustic diversity and synthesize personalized voice by generating speech using unseen target speakers’ utterances. The proposed model integrates a speaker encoder into a TTS model consisting of the FastSpeech2 acoustic model and the HiFi-GAN vocoder. The speaker encoder, based on the pre-trained RawNet3, extracts speaker-specific voice features. Furthermore, the proposed approach not only includes an English one-shot multi-speaker TTS but also introduces a Korean one-shot multi-speaker TTS. We evaluate naturalness and speaker similarity of the generated speech using objective and subjective metrics. In the subjective evaluation, the proposed Korean one-shot multi-speaker TTS obtained naturalness mean opinion score (NMOS) of 3.36 and similarity MOS (SMOS) of 3.16. The objective evaluation of the proposed English and Korean one-shot multi-speaker TTS showed a prediction MOS (P-MOS) of 2.54 and 3.74, respectively. These results indicate that the performance of our proposed model is improved over the baseline models in terms of both naturalness and speaker similarity.

    • KCI등재

      얼굴 하단 근육의 움직임을 반영한 초음파 도플러 기반 음성합성

      이기승 한국음향학회 2025 한국음향학회지 Vol.44 No.5

      The ultrasonic Doppler-based silent speech interface technology, characterized by non-contact sensing, low-cost sensors, and long-range acquisition capabilities, has shown relatively high speech recognition accuracy in previous studies focused on isolated words. In conventional ultrasonic Doppler-based silent speech interfaces, ultrasound was emitted toward the front of the lips to detect variations caused by lip shapes. However, this approach has limitations in detecting tongue movements, which are closely related to articulating phonemes. To partially overcome this limitation, this paper proposed a method that the emitted ultrasound toward the muscle area involved in tongue movement to acquire ultrasonic displacement signals, which were then used for speech synthesis. Compared to the conventional front radiation-reflection method, the proposed approach showed superior performance in objective evaluation metrics, and the synthesized speech using Whisper and gText- To-Speech (gTTS) also demonstrated excellent subjective quality.

    • KCI등재

      Mix-MaxETTS: A text-to-emotional speech synthesis model based on a deep encoder–decoder structure for the transfer of secondary emotions

      Seyyed Mahdi Hassani,Mohammad Reza Kangavari 한국전자통신연구원 2026 ETRI Journal Vol.48 No.4

      Given the importance of emotions in social interactions, emotional speech synthesishas attracted significant attention in the field of human–computer interaction. Remarkable advancements have been made in emotional textto-speech synthesis, but most previous studies have concentrated on imitatingstyles associated with a specific primary emotion, neglecting secondary emotionsthat arise from mixtures of primary emotions. Therefore, there is a needto leverage both primary and secondary emotions in speech synthesis to facilitatemore engaging, realistic, and natural interactions among artificial socialagents. To address this gap, we propose a text-to-emotional speech synthesismodel designed to generate nuanced mixtures of emotions that effectively conveysecondary emotions during interactions. By adjusting the values of eachbasic emotion, we can control the mix of emotions in the synthetic speech. Our proposed method distinguishes between primary emotions and variationsin mixed emotions while learning emotional styles. The effectiveness of theproposed framework was validated through both objective and subjectiveevaluations.

    • Emotional Communication System for Hard of Hearing

      Sanghoon Lee,Minwoo Choi,Hyejin Koo,Sangchun Park 한국차세대컴퓨팅학회 2022 한국차세대컴퓨팅학회 학술대회 Vol.2022 No.10

      There needs an opportunity for hard of hearing to express their emotions and to have a natural conversation. Communicating through text or sign languages is not easy to fully express emotions, especially in situations where conversation partners are invisible, such as on a phone call. In this paper, we design a emotional conversation system for hard of hearing based on text emotion recognition (TER), emotional speech synthesis (ESS), and speech-to-text (STT). This system allows them to convey both their opinions and feelings accurately.

    연관 검색어 추천

    이 검색어로 많이 본 자료

    활용도 높은 자료

    해외이동버튼