올바른 자세로 악기를 연주하는 것은 좋은 소리를 내는데 도움을 주고 부상을 방지하는데 도움을 줌으로 올바른 자세로 악기를 연주하는 것은 매우 중요한 문제다. 올바른 자세로 악기를 ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T15485164
서울 : 숙명여자대학교, 2020
학위논문(박사) -- 숙명여자대학교 대학원 , IT공학과 IT공학전공 , 2020
2020
한국어
004 판사항(6)
004 판사항(23)
서울
xii, 106장 : 삽화, 도표 ; 26 cm
지도교수: 박영호
AV-TEN "Audio-Visual Tensor Fusion Network"의 약어임
권말부록: 프로페셔널 피아니스트와 아마추어 피아니스트 분류 기준과 프로페셔널 피아니스트의 경력 ; C2Pap 영상 링크 ; C3Pap 세부 정보
참고문헌 수록
I804:11043-000000069026
0
상세조회0
다운로드올바른 자세로 악기를 연주하는 것은 좋은 소리를 내는데 도움을 주고 부상을 방지하는데 도움을 줌으로 올바른 자세로 악기를 연주하는 것은 매우 중요한 문제다. 올바른 자세로 악기를 ...
올바른 자세로 악기를 연주하는 것은 좋은 소리를 내는데 도움을 주고 부상을 방지하는데 도움을 줌으로 올바른 자세로 악기를 연주하는 것은 매우 중요한 문제다.
올바른 자세로 악기를 연주할 수 있도록 피아노 연주와 IT기술을 접목한 연구들이 진행되었다. 기존의 피아노 연주 자세 분석 연구는 다음과 같은 한계점이 있다. 첫째, 사람마다 연주 관련 근 골격 계 질환을 일으키는 자세가 다르기 때문에 특정 자세를 부상 자세로 정의하고 분류하는 연구가 부상 방지에 도움이 되기 어렵다는 한계점이 있다. 둘째, 기존 연구들은 연주자의 자세 정보를 얻기 위하여 연주자의 센서를 부착하여 연주자가 연주하는데 방해가 된다. 셋째, 기계 학습을 이용한 자세 분석 연구는 자세를 분류하는 특징을 수작업으로 선택해야 하고 데이터 양이 늘어날 경우 다시 특징을 선택해야 하는 번거로움이 있다. 또한, 피아노 연주 자세는 소리와 관련성이 높은데 기존의 연구는 청각 정보를 고려하지 않고 시각 정보만 활용하였다는 한계점이 있다.
본 논문에서는 시청각 정보를 활용한 딥러닝 기반의 피아노 연주 자세 분류 연구를 진행한다. 첫째, 시각 정보만 활용하거나 기계학습방법을 사용하여 피아노 연주 자세를 분류하였던 과거의 연구들과 달리 제안하는 방법은 연주 자세와 관련성이 깊은 청각 정보를 활용하여 연주 자세를 분석하는 모델을 제안한다
둘째, 본 논문에서는 시청각 정보를 하나의 데이터 구조로 표현하는 방식인 오디오-비쥬얼 결합 방법을 제안한다. 제안하는 오디오-비쥬얼 결합 방법은 이전 관련연구들과 달리 청각 정보를 컬러 스케일로 표현하고 시각정보를 흑백 스케일로 표현하고 이 둘의 색상을 섞어 하나의 데이터 구조에 표현하는 방식으로 하나의 데이터 구조에 시청각 정보를 모두 표현 할 수 있는 데이터 표현 방법이다.
셋째, 제안하는 AV-TFN과 관련 연구인 VN (Visual Network), AN (Audio Network), AVN (Audio-Visual Network)과의 성능을 비교하고 제안하는 방법의 우수함을 보였다. 제안하는 AV-TFN은 속도는 유지하면서 F1점수가 평균 약 6.16 퍼센트포인트 향상되는 결과를 보였다.
넷째, 피아노 연주 자세 데이터 셋인 C3Pap (Classic piano performance posture version amateur & pro)를 제안한다. 제안하는 C3Pap는 타 관련 연구와 비교하여 피아니스트 총 명수, 프로페셔널 피아니스트와 아마추어 피아니스트의 비율, 각 클래스 별 프로페셔널 피아니스트의 영상과 아마추어 피아니스트의 영상 비율, 곡의 총 개수, 프로페셔널 피아니스트의 영상의 경우 세계 거장 피아니스트의 영상으로만 수집 등 다양한 방면에서 우수하다.
다국어 초록 (Multilingual Abstract)
Playing the instrument in the correct position is important because the correct position helps to produce good sounds and prevents injuries. Many studies have been conducted in the field of piano performance posture recognition that combine various pi...
Playing the instrument in the correct position is important because the correct position helps to produce good sounds and prevents injuries. Many studies have been conducted in the field of piano performance posture recognition that combine various piano playing and IT techniques. However, considering that different postures cause different playing-related musculoskeletal disorders (PRMDs), studies that define and recognize specific postures as injured postures are not efficient in preventing injuries. In addition, existing studies have the following limitations: (1) existing studies that use wearable devices interfere with playing due to attachment of sensors in the body of pianist; (2) existing studies that use machine learning has a disadvantage in that a feature for classifying a pose must be manually selected and features must be extracted again as the number of data increases; (3) existing studies have only used visual information for posture classification without considering audio information.
In this paper, we propose an audio-visual tensor fusion network (simply, AV-TFN) for piano performance posture classification. Unlike existing studies that used only visual information or classifying piano postures using machine learning methods, the proposed method uses audio information to improve the accuracy in classifying the postures of professional and amateur pianists. For this, we first propose a dataset called C3Pap (Classic piano performance posture version amateur & pro). Unlike existing studies, where it was difficult to collect professional data, this study collects data from various professional and amateur pianists from the YouTube platform. In the collected dataset, the ratio of professional and amateur pianist data is even and has the advantage of having actual performance videos rather than data collected in a specific environment. Further, we propose a data structure that represents audio-visual information. The proposed data structure represents audio information in color scale and visual information on the black and white scale. We call this data structure as an audio-video tensor. Finally, we compare the proposed method with its variants: VN (Visual Network), AN (Audio Network), AVN (Audio-Visual Network). The experiment results demonstrate that AV-TFN outperforms existing studies and thus, can be effectively used in the classification of piano performance postures.
목차 (Table of Contents)