RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    검색결과 좁혀 보기

    선택해제
    • 좁혀본 항목 보기순서

      • 원문유무
      • 음성지원유무
      • 원문제공처
        펼치기
      • 등재정보
        펼치기
      • 학술지명
        펼치기
      • 주제분류
        펼치기
      • 발행연도
        펼치기
      • 작성언어

    오늘 본 자료

    • 오늘 본 자료가 없습니다.
    더보기
    • 무료
    • 기관 내 무료
    • 유료
    • KCI등재

      분산형 시스템을 적용한 음성합성에 관한 연구

      김진우,민소연,나덕수,배명진,Kim, Jin-Woo,Min, So-Yeon,Na, Deok-Su,Bae, Myung-Jin 한국음향학회 2010 韓國音響學會誌 Vol.29 No.3

      최근 광대역 무선 통신망의 보급과 소형 저장매체의 대용량화로 인하여 이동형 단말기가 주목 받고 있다. 이로 인해 이동형 단말기에 문자정보를 청취할 수 있도록 문자를 음성으로 변환해 주는 TTS(Text-to-Speech) 기능이 추가되고 있다. 사용자의 요구사항은 고음질의 음성합성이지만 고음질의 음성합성은 많은 계산량이 필요하기 때문에 낮은 성능의 이동형 단말기에 는 적합하지 않다. 본 논문에서 제안하는 분산형 음성합성기 (DTTS)는 고음질 음성합성이 가능한 코퍼스 기반 음성합성 시스템을 서버와 단말기로 나누어 구성한다. 서버 음성합성 시스템은 단말기에서 전송된 텍스트를 데이터베이스 검색 후 음성파형 연결정보를 생성하여 단말기로 전송하고, 단말기 음성합성 시스템은 서버 음성합성 시스템에서 생성된 음성파형 연결정보와 단말기에 존재하는 데이터베이스를 이용하여 간단한 연산으로 고음질 합성음을 생성할 수 있는 시스템이다. 제안하는 분산형 합성기는 단말기에서의 계산량을 줄여 저가의 CPU 사용, 전력소모의 감소, 효율적인 유지보수를 할 수 있도록 하는 장점이 있다. Recently portable terminal is received attention by wireless networks and mass capacity ROM. In this result, TTS(Text to Speech) system is inserted to portable terminal. Nevertheless high quality synthesis is difficult in portable terminal, users need high quality synthesis. In this paper, we proposed Distributed TTS (DTTS) that was composed of server and terminal. The DTTS on corpus based speech synthesis can be high quality synthesis. Synthesis system in server that generate optimized speech concatenation information after database search and transmit terminal. Synthesis system in terminal make high quality speech synthesis as low computation using transmitted speech concatenation information from server. The proposed method that can be reducing complexity, smaller power consumption and efficient maintenance.

    • KCI등재

      콜퍼스에 기반한 한국어 문장/음성변환 시스템

      김상훈,박준,이영직,Kim, Sang-hun,Park, Jun,Lee, Young-jik 한국음향학회 2001 韓國音響學會誌 Vol.20 No.3

      이 논문에서는 대용량 음성 데이터베이스를 기반으로 하는 한국어 문장/음성변환시스템의 구현에 관해 기술한다. 기존 소량의 음성데이타를 이용하여 운율조절을 통해 합성하는 방식은 여전히 기계음에 가까운 합성음을 생성하고 있다. 이러한 문제점을 해결하기 위해 본 논문에서는 대용량 음성 데이터베이스를 기반으로 하여 운율처리없이 합성단위 선정/연결에 의해 합성음질을 향상시키고자 한다. 대용량 음성 데이터베이스는 다양한 운율변화를 포함하도록 문장단위를 녹음하며 이로부터 복수개의 합성단위를 추출, 구축한다. 합성단위는 음성인식기를 훈련, 자동으로 음소분할하여 생성하며, 래링고그라프 신호를 이용하여 정교한 피치를 추출한다. 끊어 읽기는 휴지길이에 따라 4단계로 설정하고 끊어읽기 추정은 품사열의 통계정보를 이용한다. 합성단위 선정은 운율/스펙트럼 파라미터를 이용하여 비터비 탐색을 수행하게 되며 유클리디언 누적거리가 최소인 합성단위열을 선정/연결하여 합성한다. 또한 이 논문에서는 고품질 음성합성을 위해 특정 서비스 영역에 더욱 자연스러운 합성음을 생성할 수 있는 영역의존 음성합성용 데이터베이스를 제안한다. 구현된 합성시스템은 주관적 평가방법으로 명료도와 자연성을 평가하였고 그 결과 대용량 음성 데이터베이스를 기반으로한 합성방식의 성능이 기존 반음절단위를 사용한 합성방식보다 더 나은 성능을 보임을 알 수 있었다. this paper describes a baseline for an implementation of a corpus-based Korean TTS system. The conventional TTS systems using small-sized speech still generate machine-like synthetic speech. To overcome this problem we introduce the corpus-based TTS system which enables to generate natural synthetic speech without prosodic modifications. The corpus should be composed of a natural prosody of source speech and multiple instances of synthesis units. To make a phone level synthesis unit, we train a speech recognizer with the target speech, and then perform an automatic phoneme segmentation. We also detect the fine pitch period using Laryngo graph signals, which is used for prosodic feature extraction. For break strength allocation, 4 levels of break indices are decided as pause length and also attached to phones to reflect prosodic variations in phrase boundaries. To predict the break strength on texts, we utilize the statistical information of POS (Part-of-Speech) sequences. The best triphone sequences are selected by Viterbi search considering the minimization of accumulative Euclidean distance of concatenating distortion. To get high quality synthesis speech applicable to commercial purpose, we introduce a domain specific database. By adding domain specific database to general domain database, we can greatly improve the quality of synthetic speech on specific domain. From the subjective evaluation, the new Korean corpus-based TTS system shows better naturalness than the conventional demisyllable-based one.

    • KCI등재

      장애인을 위한 음성 인터페이스의 UI/UX: 음성 명령어 인식기와 음성 합성기를 대상으로

      홍기형,이희연,김선희,조남현,김지환,정민화 에스케이텔레콤 (주) 2013 Telecommunications Review Vol.23 No.2

      본 논문은 장애인과 함께 사는 사회를 추구하는 QoLT 연구개발과제 가운데 음성인식 및 음성합성 기반 장애인용 음성 인터페이스에 대한 사용자 중심의 사용성 평가를 통하여 장애인용 음성 인터페이스의 UI/UX 문제를 고찰하는 것을 그 목적으로 한다. 장애인용 음성 인터페이스로는 마비말장애인용 음성 명령어 인식기와 시각장애인용 음성합성기를 대상으로 하여 그에 대한 사용성 평가를 중심으로 살펴보았다. 마비말장애인용 음성 명령어 인식기의 경우는 총 2차에 걸쳐 53명이 참여한 사용성 평가를 통하여 UI 적절성 및 기능적 측면을 포함하는 요구사항을 분석하였다. 시각장애인용 음성합성기의 경우는 컴퓨터에 저장된 합성음 파일을 이용하여 기존의 고속합성음과 제안된 합성음을 비교하고 여성 및 남성 합성음의 음색별 선호도 평가를 시행하였다. 본 연구의 결과, 음성 인터페이스 개발 이전과 개발 도중에 사용자 그룹인 마비말장애인과 시각장애인들을 대상으로 사용자의 요구 사항 및 사용성 평가를 시행하여 각각의 프로토타이프 개발에 반영하고, 음성 인터페이스의 보급과 필요성에 관한 긍정적인 사용자의 반응을 이끌어 낼 수 있었다.

    • 다양한 발성에 따른 다중음성 합성 시스템

      박현영 ( Hyun Young Park ),김명 ( Myoung Kim ),배명진 ( Myoung Jin Bae ) 한국감성과학회 2003 한국감성과학회 추계학술대회 Vol.2003 No.-

      음성 합성이란 기계적인 장치나 전지회로 또는 컴퓨터 모의를 이용하여 자동으로 음성파형을 생성해 내는 것으로 정의한다. 음성 합성에 대한 연구는 다른 음성에 관련된 기술들보다 가장 먼저 연구된 기술이다. 음성 합성기는 PC의 보급이 확대되고 통신 시장이 컴짐에 따라 그 응용 분야가 점차 확대되어 가고 다양한 방식의 음성 합성 기법에 관한 연구가 이루어지고 있다. 일반적으로 자연스러운 대화를 할 때나 글을 읽을 때의 음성에는 퍼지, 지속시간, 에너지 등의 운율 정보가 포함되어 있다. 따라서, 문장을 합성하는 경우 운율정보를 합성음에 반영하면 보다 명확한 의미 전달과 다양한 발성변환이 가능해 진다. 본 논문에서는 시간영역에서 PSOLA 합성방식에 의한 피치 변경과 지속시간 변경을 이용하여 다양한 발성변환에 따른 다중음성 합성기를 구현하였다.

    • KCI등재

      시각장애인용 음성합성기에 대한 사용자 요구분석

      이희연,홍기형 이화여자대학교 특수교육연구소 2012 특수교육 Vol.11 No.2

      본 연구의 목적은 시각장애인용 보조기기인 스크린리더에 사용되는 음성합성기에 대한 사용자의 요구분석 사항을 파악하여 고품질의 합성음을 개발하여 멀티미디어 환경 내에서 시각 장애인들의 지식정보 접근성을 높이고, 스크린리더 등의 전용 프로그램의 기능을 보완 ·강화하여 교육 및 사회 참여의 기회를 확대하여 전반적인 삶의 질을 향상하는데 있다. 컴퓨터와 스크린 리더의 사용에 익숙한 총 다섯 명의 참가자가 본 연구를 위한 핵심집단 면담연구에 참여했으며, 참가자의 요구분석 내용들은 상향식 접근방법(bottom-up approach)을 통하여 분석되었다. 시각장애인들의 음성합성기에 대한 요구분석들을 분석한 결과, (1) 청취 시의 피로도 개선, (2) 합성음의 명료도 개선, (3) 자연스러운 합성음에 대한 요구, (4) 남성 음색의 합성음에 대한 요구, (5) 잡음 발생 최소화, (6) 포네틱(phonetic) 기능에 대한 요구, (7) 정보유형에 따른 차별화된 속도 제공 등의 요구를 파악할 수 있었다. 이러한 결과에 근거하여 다양한 남성 음색의 명료하고 자연스러운 음성합성기를 개발하고, 이에 대한 시각장애인들의 장기간에 걸친 사용성 및 만족도를 평가하는 후속 연구가 요구된다. The purpose of this study was to examine usability and needs for speech synthesizer used in a screen reader which is an assistive technology device for the blind in order to (1) improve web accessibility in multimedia environment, (2) enhance the quality of a screen reader, (3) expand educational and vocational opportunities, and (4) improve the overall quality of life for the blind. Participants were five individuals (three men and two women) who were blind between the ages of thirty five to forty nine years old, and all of them were good at using computers, keyboards, and screen reader programs. After each participant was participated in the nonsense-word dictation task and the sentence listening task presented at two different speeds, they were engaged in a focus group interview to examine their preferences, needs, interests, and other opinions regarding the synthesized speech. The results of this study were analyzed into following categories using a bottom-up approach. Participants’ needs for the speech synthesizer were (1) reduction of fatigue while listening synthesized speech for a long time, (2) improvement of the intelligibility of the synthesized speech, (3) use of natural and comfortable synthesized speech, (4) use of low tone of voice (male voice tone), (5) removal of background noises or high-frequency noises, (6) addition of the phonetic reading function, and (7) differentiated speed of synthesized speech based on the type of information. The first priority of the speech synthesizer users’ was to minimize fatigue level while listening a synthesized speech. Further research needs to be directed to examine usability and needs for the updated speech synthesizer through a long-term usability testing.

    • DCGAN 의 잠재 벡터 보간을 활용한 두 음성 합성 방법

      허찬영,정재희 한국차세대컴퓨팅학회 2023 한국차세대컴퓨팅학회 학술대회 Vol.2023 No.06

      기계 학습 및 딥러닝 기술의 발전은 문학 분야를 비롯한 다양한 예술 분야에서 인공지능이 그림을 그리고 소설을 쓰거나 음악을 작곡, 작사하는 것과 같이 큰 영향력을 끼치고 있다. 이 중 인공지능이 음악을 작곡, 작사하는 음성을 생성하는 분야에서도 이미지 생성에 특화된 GANs(Generative Adversarial Nets) 모델을 사용하여 음성을 생성하는 연구를 적용할 수 있다. 하지만 음성 데이터 자체로 학습하여 음성을 생성하는 데에는 GANs를 사용할 경우 적절한 음성 생성의 결과를 얻지 못한다. 따라서 음성을 이미지로 변환하여 GANs을 학습한 후, 이미지를 생성하여 이를 다시 음성으로 생성하는 방법으로 음성 생성을 할 수 있다. 본 연구에서는 CNN(Convolution Neural Network) 기반의 GANs 모델인 DCGAN(Deep Convolutional Generative Adversarial Network) 모델을 활용하여, 두 개의 생성된 음성 이미지에서 추출된 잠재 벡터 z들의 보간의 정도에 따라 생성된 이미지가 부드럽게 변하는 특징을 적용하여 음성 합성 방법을 제안한다. 두 개의 서로 다른 음성 포맷인 midi 파일과 wav 파일을 각각 이미지로 변환 후 모델을 학습시켰다. 두 포맷 모두 두개의 음성 이미지의 잠재 벡터의 보간 정도에 따라 생성된 이미지가 부드럽게 변환되었고, 각 보간 값의 정도에 따라 생성된 이미지들을 다시 음성으로 변환시켜 적절히 합성된 음성을 확인할 수 있었다.

    • KCI등재

      음성기술의 발전과 국어음성학의 역할 -인간과 비인간, 인간다움과 인간스러움의 경계 이동-

      송민규,김숙정 겨레어문학회 2026 겨레어문학 Vol.76 No.-

      본고의 목적은 음성인식과 음성합성 기술의 발전 과정 속에서 인간과 비인간의 경계가 재편되는 양상을 고찰하는 데 있다. 음성기술의 발전사는 공학적 진보의 역사이면서 동시에 인간성을 재정의하게 만드는 존재론적 전환의 역사이다. 먼저 인간의 의사소통 경로가 기술적으로 어떻게 모사되는지를 통합 언어 연쇄의 관점에서 정리하고, 이어 음성인식 기술이 인간 청각과 해석 행위를 어떻게 알고리즘화해 왔는지를 검토하였다. 다음으로 음성합성 기술의 계보를 통해 인간 음성이 비신체화‧정보화되는 과정을 논의하고, 이를 통해 음성기술의 변화가 인간과 기계의 새로운 상호작용 질서를 구성하는 과정임을 확인하였다. 나아가 본고는 음성기술과 관련된 사회적·윤리적 쟁점을 분석하고, 포용성, 현실성, 지역성, 보안성의 관점에서 향후 국어 음성학이 수행해야 할 과제를 제안하였다. The purpose of this study is to examine the reconfiguration of boundaries between humans and non-humans by tracing the evolution of speech recognition and synthesis technologies. The history of speech technology is not merely a record of engineering advancement but a trajectory of ontological shifts that compel a redefinition of humanity. First, this paper delineates the technical simulation of human communication channels from the perspective of an "integrated speech path." Subsequently, it investigates the process through which speech recognition technology has algorithmized human audition and interpretive practices. By exploring the genealogy of speech synthesis, the study further discusses the disembodiment and informatization of the human voice, ultimately asserting that the evolution of speech technology establishes a new order of interaction between humans and machines. Finally, this paper analyzes the social and ethical controversies surrounding speech technology and proposes future research directions for Korean phonetics, emphasizing the dimensions of inclusivity, realism, locality, and security.

    • KCI등재

      토픽모델링과 네트워크 분석에 기반한 AI 음성기술 연구 동향 분석

      장관종 국제차세대융합기술학회 2024 차세대융합기술학회논문지 Vol.8 No.9

      본 연구에서는 토픽모델링과 네트워크 분석을 활용하여 AI 음성기술의 연구 동향을 WoS에 등재된 한국저자 논문 1,530편을 대상으로 3차 시기로 나누어 AI 음성기술 연구의 지적 네트워크과 주요 연구 토픽을 분석했다. 초기 연구(2011-2015)는 HMM과 같은 머신러닝 모델 및 초기 딥 러닝 기술을 사용하여 음성인식, 소음 감소및 신호 처리 개선에 중점을 둔 강력한 음성인식, 청각 처리 및 화자 적응이었다. 중기(2016~2020)는 딥 러닝과머신러닝 기술을 적용하여 음성 및 언어 처리 분야에서 상당한 발전을 가져왔고 감정 인식, 시끄러운 환경에서의향상된 음성인식 및 의료 분야와 같은 응용 분야에 중점을 두고 있었다. 최근 연구(2021~2024)는 자연어 처리를위한 Transformers 및 BERT를 포함한 정교한 AI 모델을 사용하여 음성 및 감정 인식이 지속적으로 발전하고있고, 맞춤형 음성합성, 달팽이관 이식과 같은 보조 기술의 적용에도 중점은 두고 있다. 본 연구에서 한국은 2011 년부터 2024년까지 AI 음성처리 기술 분야의 연구 동향을 분석한 결과는 상당한 기술 발전과 적용 확대를 경험했다는 결론이 나왔다. 이러한 발전은 지속적인 혁신과 새로운 과제에 대한 적응을 통해 다양한 부문에서 AI 음성처리 기술의 영향력과 중요성이 커지고 있음을 반영할 수 있다. In this study, using topic modeling and network analysis, research trends in AI voice technology were divided into three periods targeting 1,530 papers by Korean authors registered in WoS to identify the intellectual network and major research topics of AI voice technology research. was analyzed. Early research (2011-2015) was in robust speech recognition, auditory processing and speaker adaptation, focusing on improving speech recognition, noise reduction and signal processing using machine learning models such as the HMM and early deep learning techniques. The mid-term (2016-2020) brought significant advances in the field of speech and language processing by applying deep learning and machine learning technologies, focusing on application areas such as emotion recognition, improved speech recognition in noisy environments, and the medical field. Recent research (2021-2024) continues to advance speech and emotion recognition using sophisticated AI models, including Transformers and BERT for natural language processing, and also focuses on the application of assistive technologies such as personalized speech synthesis and cochlear implants. there is. In this study, the results of analyzing research trends in the field of AI voice processing technology from 2011 to 2024 concluded that Korea has experienced significant technological development and expansion of application. These developments may reflect the growing influence and importance of AI voice processing technology in various sectors through continuous innovation and adaptation to new challenges.

    • KCI등재

      TTS 음성에서 말속도, 메시지 유형 및 청취자 연령이 음성 호감도에 미치는 영향

      하나경,진유나,박윤지,심은진,이영미 한국언어치료학회 2025 언어치료연구 Vol.34 No.2

      목적: 본 연구는 TTS 말속도, 메시지 유형, 청자 연령이 음성 호감도에 미치는 영향을 분석하였다. 방법: TTS 음성 목록 중 두 개의 여성 화자(20~30대) 음성 선정한 뒤, 세 가지 말속도(느림, 보통, 빠름)와 두 가지 메시지 유형(이성적, 감성적)을 조합하여 총 12개의 음성을 생성하였다. 총 40명의 참가자가 연구에 참여하였으며, 청년층(n=20), 노년층(n=20)으로 구분되었다. 각 음성에 대한 호감도 평가 후, 연령대별 음성 호감도 차이를 분석하였다. 결과: TTS 말속도는 음성 호감도에 유의미한 영향을 미쳐 보통 속도보다 느리거나 빠를 때 더 높은 호감도를 보였다. 감성적 메시지는 이성적 메시지보다, 노년층은 청년층보다 호감도가 높았다. 말속도와 메시지 유형 간 상호작용에서 감성적 메시지는 느린 속도에서, 이성적 메시지는 빠른 속도에서 더 높은 호감도를 보였으며, 삼차 상호작용은 유의하지 않았다. 결론: 본 연구는 TTS 음성의 말속도가 음성 호감도에 미치는 영향을 분석하고, 메시지 유형 및 청자의 연령대와의 상호작용 효과를 검증하였다. 연구 결과, TTS 말속도는 음성 호감도에 유의미한 영향을 미쳤으며, 느린 속도와 보통 속도, 보통 속도와 빠른 속도 간 유의미한 차이가 관찰되었다. 감성적 메시지는 느린 속도에서 가장 높은 호감도를 보였고, 이성적 메시지는 빠른 속도에서 가장 긍정적인 평가를 받았다. 또한, 노년층은 청년층보다 전반적으로 더 높은 호감도를 나타냈다. 이러한 결과는 TTS 음성 설계에서 메시지 유형과 말속도를 맥락에 맞게 조정함으로써 사용자 경험을 최적화할 수 있는 가능성을 제시하며, 보완대체의사소통(AAC) 시스템과 같은 응용 분야에서의 활용을 시사한다. Purpose: This study investigated the effects of TTS (text-to-speech) speech rate, message type, and listener age on the perception of voice attractiveness. Methods: Two TTS-generated female voices (speakers in their 20~30s) were manipulated to produce speech at three different rates (slow, moderate, fast) and deliver two distinctive message types (emotional, rational). A total of 40 participants, comprising younger adults (n=20) and older adults (n=20), evaluated the attractiveness of these voices using a 5-point Likert scale. Results: TTS speech rate had a statistically significant effect on voice attractiveness, with slower or faster rates showing higher likability than moderate speeds. Emotional messages were rated more favorably than rational ones, and older adults provided higher overall ratings compared to younger adults. In terms of interaction effects, emotional messages were rated higher at slower rates, while rational messages were rated higher at faster rates. Conclusions: TTS speech rate was found to be a critical factor in voice attractiveness. Differences emerged between slow and normal speeds, as well as between normal and fast speeds. Emotional messages were most appealing at a slow rate, whereas rational messages were most favorably received at a fast rate. Older adults, overall, provided higher ratings than younger adults. These findings suggest that tailoring TTS speech rate and message type to specific contexts can enhance user experience, offering practical insights for AAC systems and related applications.

    • KCI등재

      스펙트럼 형태 불변 실시간 음성 변환 시스템

      김원구(Weon-Goo Kim) 한국지능시스템학회 2005 한국지능시스템학회논문지 Vol.15 No.1

      본 논문에서는 음성의 스펙트럼 형태는 유지하면서 음성을 기계적인 음성으로 변환시키기는 실시간 음성 변환 방법을 제안하였다. 이러한 목적을 위하여 LPC 분석 및 합성 방법을 사용하여 변환된 음성의 스펙트럼은 유지하였고 합성된 음성의 피치는 자유롭게 변경되도록 하였다. 제안된 방법에서는 변환된 음성이 보다 자연스럽게 들리게 하기 위하여 여기 신호 발생기에 이득 정합 방법을 적용하였다. 제안된 방법의 성능을 평가하기 위하여 음성 변환 실험을 수행하였다. 실험 결과에서 원 음성 신호는 원 화자의 신원을 알기가 어려운 기계적인 음성 신호로 바뀌는 것을 알 수 있었고 피치의 심한 변화에도 변환된 음성의 의미는 정확히 전달될 수 있었다. 제안된 시스템은 시스템의 실시간으로 구현될 수 있는지 확인하기 위하여 TI TMS320C6711DSK 보드를 사용하여 구현되었다 In this paper, the spectral shape invariant real-time voice change method is proposed to change one's voice to mechanical voice. For this purpose, LPC analysis and synthesis is used to maintain the spectraum of voice and the pitch of synthesis speech can be changed freely. In the proposed method, gain matching method is applied to excitation signal generator to make the changed voice natural to hear. In order to evaluate the performance of the proposed method, voice change experiments were conducted. Experimental results showed that original speech signal is changed to the mechanical voice signal in which context of the speaker's voice is conveyed correctly in spite of drastic change of pitch. The system is implemented using TI TMS320C6711DSK board to verify the system runs in real time.

    연관 검색어 추천

    이 검색어로 많이 본 자료

    활용도 높은 자료

    해외이동버튼