RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    AV-TFN : 피아노 연주 자세 분류를 위한 오디오-비쥬얼 결합 텐서 네트워크 = AV-TFN : Audio-Visual Tensor Fusion Network for piano posture classification

    한글로보기

    https://www.riss.kr/link?id=T15485164

    • 저자
    • 발행사항

      서울 : 숙명여자대학교, 2020

    • 학위논문사항
    • 발행연도

      2020

    • 작성언어

      한국어

    • KDC

      004 판사항(6)

    • DDC

      004 판사항(23)

    • 발행국(도시)

      서울

    • 형태사항

      xii, 106장 : 삽화, 도표 ; 26 cm

    • 일반주기명

      지도교수: 박영호
      AV-TEN "Audio-Visual Tensor Fusion Network"의 약어임
      권말부록: 프로페셔널 피아니스트와 아마추어 피아니스트 분류 기준과 프로페셔널 피아니스트의 경력 ; C2Pap 영상 링크 ; C3Pap 세부 정보
      참고문헌 수록

    • UCI식별코드

      I804:11043-000000069026

    • 소장기관
      • 국립중앙도서관 국립중앙도서관 우편복사 서비스
      • 숙명여자대학교 도서관 소장기관정보
    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    올바른 자세로 악기를 연주하는 것은 좋은 소리를 내는데 도움을 주고 부상을 방지하는데 도움을 줌으로 올바른 자세로 악기를 연주하는 것은 매우 중요한 문제다.
    올바른 자세로 악기를 연주할 수 있도록 피아노 연주와 IT기술을 접목한 연구들이 진행되었다. 기존의 피아노 연주 자세 분석 연구는 다음과 같은 한계점이 있다. 첫째, 사람마다 연주 관련 근 골격 계 질환을 일으키는 자세가 다르기 때문에 특정 자세를 부상 자세로 정의하고 분류하는 연구가 부상 방지에 도움이 되기 어렵다는 한계점이 있다. 둘째, 기존 연구들은 연주자의 자세 정보를 얻기 위하여 연주자의 센서를 부착하여 연주자가 연주하는데 방해가 된다. 셋째, 기계 학습을 이용한 자세 분석 연구는 자세를 분류하는 특징을 수작업으로 선택해야 하고 데이터 양이 늘어날 경우 다시 특징을 선택해야 하는 번거로움이 있다. 또한, 피아노 연주 자세는 소리와 관련성이 높은데 기존의 연구는 청각 정보를 고려하지 않고 시각 정보만 활용하였다는 한계점이 있다.
    본 논문에서는 시청각 정보를 활용한 딥러닝 기반의 피아노 연주 자세 분류 연구를 진행한다. 첫째, 시각 정보만 활용하거나 기계학습방법을 사용하여 피아노 연주 자세를 분류하였던 과거의 연구들과 달리 제안하는 방법은 연주 자세와 관련성이 깊은 청각 정보를 활용하여 연주 자세를 분석하는 모델을 제안한다
    둘째, 본 논문에서는 시청각 정보를 하나의 데이터 구조로 표현하는 방식인 오디오-비쥬얼 결합 방법을 제안한다. 제안하는 오디오-비쥬얼 결합 방법은 이전 관련연구들과 달리 청각 정보를 컬러 스케일로 표현하고 시각정보를 흑백 스케일로 표현하고 이 둘의 색상을 섞어 하나의 데이터 구조에 표현하는 방식으로 하나의 데이터 구조에 시청각 정보를 모두 표현 할 수 있는 데이터 표현 방법이다.
    셋째, 제안하는 AV-TFN과 관련 연구인 VN (Visual Network), AN (Audio Network), AVN (Audio-Visual Network)과의 성능을 비교하고 제안하는 방법의 우수함을 보였다. 제안하는 AV-TFN은 속도는 유지하면서 F1점수가 평균 약 6.16 퍼센트포인트 향상되는 결과를 보였다.
    넷째, 피아노 연주 자세 데이터 셋인 C3Pap (Classic piano performance posture version amateur & pro)를 제안한다. 제안하는 C3Pap는 타 관련 연구와 비교하여 피아니스트 총 명수, 프로페셔널 피아니스트와 아마추어 피아니스트의 비율, 각 클래스 별 프로페셔널 피아니스트의 영상과 아마추어 피아니스트의 영상 비율, 곡의 총 개수, 프로페셔널 피아니스트의 영상의 경우 세계 거장 피아니스트의 영상으로만 수집 등 다양한 방면에서 우수하다.
    번역하기

    올바른 자세로 악기를 연주하는 것은 좋은 소리를 내는데 도움을 주고 부상을 방지하는데 도움을 줌으로 올바른 자세로 악기를 연주하는 것은 매우 중요한 문제다. 올바른 자세로 악기를 ...

    올바른 자세로 악기를 연주하는 것은 좋은 소리를 내는데 도움을 주고 부상을 방지하는데 도움을 줌으로 올바른 자세로 악기를 연주하는 것은 매우 중요한 문제다.
    올바른 자세로 악기를 연주할 수 있도록 피아노 연주와 IT기술을 접목한 연구들이 진행되었다. 기존의 피아노 연주 자세 분석 연구는 다음과 같은 한계점이 있다. 첫째, 사람마다 연주 관련 근 골격 계 질환을 일으키는 자세가 다르기 때문에 특정 자세를 부상 자세로 정의하고 분류하는 연구가 부상 방지에 도움이 되기 어렵다는 한계점이 있다. 둘째, 기존 연구들은 연주자의 자세 정보를 얻기 위하여 연주자의 센서를 부착하여 연주자가 연주하는데 방해가 된다. 셋째, 기계 학습을 이용한 자세 분석 연구는 자세를 분류하는 특징을 수작업으로 선택해야 하고 데이터 양이 늘어날 경우 다시 특징을 선택해야 하는 번거로움이 있다. 또한, 피아노 연주 자세는 소리와 관련성이 높은데 기존의 연구는 청각 정보를 고려하지 않고 시각 정보만 활용하였다는 한계점이 있다.
    본 논문에서는 시청각 정보를 활용한 딥러닝 기반의 피아노 연주 자세 분류 연구를 진행한다. 첫째, 시각 정보만 활용하거나 기계학습방법을 사용하여 피아노 연주 자세를 분류하였던 과거의 연구들과 달리 제안하는 방법은 연주 자세와 관련성이 깊은 청각 정보를 활용하여 연주 자세를 분석하는 모델을 제안한다
    둘째, 본 논문에서는 시청각 정보를 하나의 데이터 구조로 표현하는 방식인 오디오-비쥬얼 결합 방법을 제안한다. 제안하는 오디오-비쥬얼 결합 방법은 이전 관련연구들과 달리 청각 정보를 컬러 스케일로 표현하고 시각정보를 흑백 스케일로 표현하고 이 둘의 색상을 섞어 하나의 데이터 구조에 표현하는 방식으로 하나의 데이터 구조에 시청각 정보를 모두 표현 할 수 있는 데이터 표현 방법이다.
    셋째, 제안하는 AV-TFN과 관련 연구인 VN (Visual Network), AN (Audio Network), AVN (Audio-Visual Network)과의 성능을 비교하고 제안하는 방법의 우수함을 보였다. 제안하는 AV-TFN은 속도는 유지하면서 F1점수가 평균 약 6.16 퍼센트포인트 향상되는 결과를 보였다.
    넷째, 피아노 연주 자세 데이터 셋인 C3Pap (Classic piano performance posture version amateur & pro)를 제안한다. 제안하는 C3Pap는 타 관련 연구와 비교하여 피아니스트 총 명수, 프로페셔널 피아니스트와 아마추어 피아니스트의 비율, 각 클래스 별 프로페셔널 피아니스트의 영상과 아마추어 피아니스트의 영상 비율, 곡의 총 개수, 프로페셔널 피아니스트의 영상의 경우 세계 거장 피아니스트의 영상으로만 수집 등 다양한 방면에서 우수하다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Playing the instrument in the correct position is important because the correct position helps to produce good sounds and prevents injuries. Many studies have been conducted in the field of piano performance posture recognition that combine various piano playing and IT techniques. However, considering that different postures cause different playing-related musculoskeletal disorders (PRMDs), studies that define and recognize specific postures as injured postures are not efficient in preventing injuries. In addition, existing studies have the following limitations: (1) existing studies that use wearable devices interfere with playing due to attachment of sensors in the body of pianist; (2) existing studies that use machine learning has a disadvantage in that a feature for classifying a pose must be manually selected and features must be extracted again as the number of data increases; (3) existing studies have only used visual information for posture classification without considering audio information.
    In this paper, we propose an audio-visual tensor fusion network (simply, AV-TFN) for piano performance posture classification. Unlike existing studies that used only visual information or classifying piano postures using machine learning methods, the proposed method uses audio information to improve the accuracy in classifying the postures of professional and amateur pianists. For this, we first propose a dataset called C3Pap (Classic piano performance posture version amateur & pro). Unlike existing studies, where it was difficult to collect professional data, this study collects data from various professional and amateur pianists from the YouTube platform. In the collected dataset, the ratio of professional and amateur pianist data is even and has the advantage of having actual performance videos rather than data collected in a specific environment. Further, we propose a data structure that represents audio-visual information. The proposed data structure represents audio information in color scale and visual information on the black and white scale. We call this data structure as an audio-video tensor. Finally, we compare the proposed method with its variants: VN (Visual Network), AN (Audio Network), AVN (Audio-Visual Network). The experiment results demonstrate that AV-TFN outperforms existing studies and thus, can be effectively used in the classification of piano performance postures.
    번역하기

    Playing the instrument in the correct position is important because the correct position helps to produce good sounds and prevents injuries. Many studies have been conducted in the field of piano performance posture recognition that combine various pi...

    Playing the instrument in the correct position is important because the correct position helps to produce good sounds and prevents injuries. Many studies have been conducted in the field of piano performance posture recognition that combine various piano playing and IT techniques. However, considering that different postures cause different playing-related musculoskeletal disorders (PRMDs), studies that define and recognize specific postures as injured postures are not efficient in preventing injuries. In addition, existing studies have the following limitations: (1) existing studies that use wearable devices interfere with playing due to attachment of sensors in the body of pianist; (2) existing studies that use machine learning has a disadvantage in that a feature for classifying a pose must be manually selected and features must be extracted again as the number of data increases; (3) existing studies have only used visual information for posture classification without considering audio information.
    In this paper, we propose an audio-visual tensor fusion network (simply, AV-TFN) for piano performance posture classification. Unlike existing studies that used only visual information or classifying piano postures using machine learning methods, the proposed method uses audio information to improve the accuracy in classifying the postures of professional and amateur pianists. For this, we first propose a dataset called C3Pap (Classic piano performance posture version amateur & pro). Unlike existing studies, where it was difficult to collect professional data, this study collects data from various professional and amateur pianists from the YouTube platform. In the collected dataset, the ratio of professional and amateur pianist data is even and has the advantage of having actual performance videos rather than data collected in a specific environment. Further, we propose a data structure that represents audio-visual information. The proposed data structure represents audio information in color scale and visual information on the black and white scale. We call this data structure as an audio-video tensor. Finally, we compare the proposed method with its variants: VN (Visual Network), AN (Audio Network), AVN (Audio-Visual Network). The experiment results demonstrate that AV-TFN outperforms existing studies and thus, can be effectively used in the classification of piano performance postures.

    더보기

    목차 (Table of Contents)

    • Ⅰ. 서론 = 1
    • Ⅱ. 배경지식 = 4
    • 1. 피아노 연주 테크닉과 자세 = 4
    • 1.1. 피아노 연주 테크닉 = 4
    • Ⅰ. 서론 = 1
    • Ⅱ. 배경지식 = 4
    • 1. 피아노 연주 테크닉과 자세 = 4
    • 1.1. 피아노 연주 테크닉 = 4
    • 1.2. 자세 분류 = 7
    • 1.2.1. 피아노 연주 자세 = 7
    • 1.2.2. 시각 정보만 사용할 경우의 문제점 = 7
    • 1.2.3. 청각 정보를 이용한 자세 분류 = 9
    • 2. 자세 분류를 위한 딥러닝 알고리즘 = 9
    • 2.1. Convolutional Neural Network (CNN) = 10
    • 2.2. Recurrent Neural Network (RNN) = 10
    • 2.3. Long-Short Term Memory (LSTM) = 11
    • 2.4. Gated Recurrent Units (GRU) = 13
    • Ⅲ. 관련 연구 = 14
    • 1. 피아노 연주 자세 분류 연구 = 14
    • 1.1. 자세 분류 연구 = 14
    • 1.2. 프로페셔널과 아마추어 피아니스트의 움직임 분석 연구 = 16
    • 1.3. 기존 방법들의 문제점 = 17
    • 2. 다양한 도메인의 자세분류연구 = 17
    • 2.1. 시각 정보를 이용한 자세 분류 연구 = 17
    • 2.2. 시청각 정보 기반의 자세 분류 연구 = 18
    • 2.3. 기존 방법들의 문제점 = 19
    • Ⅳ. 제안하는 오디오-비쥬얼 결합 텐서 네트워크 = 20
    • 1. 전체 프로세스 = 20
    • 2. 기존 방법과의 차이점 = 24
    • 3. 데이터 수집 = 28
    • 4. 특징 추출 = 32
    • 5. 데이터 정규화 = 34
    • 5.1. Procrustes Transformation = 34
    • 5.2. 사 분위 방법을 이용한 스켈레톤 인식률 저하 문제 보완 = 36
    • 6. 오디오-비쥬얼 결합 텐서 = 40
    • 6.1. 비디오 텐서 = 40
    • 6.2. 오디오 텐서 = 43
    • 6.3. 오디오-비쥬얼 결합 텐서 = 46
    • 7. 모델 훈련 = 51
    • 8. 모델 최적화 = 52
    • Ⅴ. 실험 = 54
    • 1. 실험 환경 = 54
    • 1.1. 컴퓨터 사양 = 55
    • 1.2. 데이터 셋 = 55
    • 1.3. 모델 구조 = 55
    • 2. 실험 = 59
    • 2.1. 시각 정보를 이용한 피아노 연주 자세 분류 모델 = 59
    • 2.1.1. 실험 환경 = 59
    • 2.1.2. 실험결과분석 = 61
    • 2.2. 청각 정보를 이용한 피아노 연주 자세 분류 모델 = 65
    • 2.2.1. 실험 환경 = 65
    • 2.2.2. 실험결과분석 = 66
    • 2.3. 시청각 정보를 이용한 피아노 연주 자세 분류 모델 = 68
    • 2.3.1. Audio-Visual Network (AVN) = 68
    • 2.3.1.1. 실험 환경 = 69
    • 2.3.1.2. 실험결과분석 = 70
    • 2.3.2. Audio-Visual Tensor Fusion Network (AV-TFN) = 72
    • 2.3.2.1. 실험 환경 = 72
    • 2.3.2.2. 실험결과분석 = 72
    • 2.4. VN, AN, AVN, AV-TFN 성능 분석 = 80
    • 2.4.1. F1점수 = 79
    • 2.4.2. 속도 = 80
    • Ⅵ. 결론 = 84
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼