본 연구는 홀로그래픽 광학 소자(Holographic Optical Element, HOE)를 사용한 차세대 디스플레이에서의 실감형 콘텐츠 조작을 위한 비전 트랜스포머(Vision Transformer, ViT) 기반의 손 제스처 인식 시스템...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=A109951923
2025
Korean
홀로그래픽 광학 소자 ; 비전트랜스포머 ; 합성곱 신경망 ; 컴퓨터 비전 ; 손 인식 ; Holographic Optical Element ; Vision Transformer ; CNN ; Hand Tracking
KCI등재
학술저널
727-734(8쪽)
0
상세조회0
다운로드본 연구는 홀로그래픽 광학 소자(Holographic Optical Element, HOE)를 사용한 차세대 디스플레이에서의 실감형 콘텐츠 조작을 위한 비전 트랜스포머(Vision Transformer, ViT) 기반의 손 제스처 인식 시스템...
본 연구는 홀로그래픽 광학 소자(Holographic Optical Element, HOE)를 사용한 차세대 디스플레이에서의 실감형 콘텐츠 조작을 위한 비전 트랜스포머(Vision Transformer, ViT) 기반의 손 제스처 인식 시스템 사용 가능성을 탐색한다. HOE는 특정 파장의 빛의 경로를 바꾸는 동작을 수행하도록 설계된 광학 소자이다. HOE 필름을 사용하여 제작한 실감 디스플레이와 사용자간 인터랙션을 위해서는 기존의 사용자와 디스플레이간 인터랙션에 사용하던 물리적 버튼이나 음성 인식을 통한 조작보다는 비접촉식 조작 시스템을 사용한 4D 조작을 통해 더 직관적인 사용자 경험을 제공하는 것이 중요하다. 이를 위해 본 논문에서는 ViT 기반의 손 제스처 인식 시스템을 제안한다. ViT는 기존의 자연어 처리 문제에서 사용되던 트랜스포머 구조를 이미지 처리 영역으로 확장한 것이다. ViT는 이미지 분류 문제에 있어 주로 사용되는 CNN과는 필터를 사용하지 않고 전역적인 Self-Attention을 통해 입력 전체의 상호작용 및 관계를 학습한다는 점에서 차별성을 가진다. 이 연구에서는 Intel RealSense D455 센서를 사용해 RGB값과 Depth값을 동시에 추출하여 생성한 커스텀 데이터셋을 이용해 학습한 ViT 모델과 CNN 모델을 동일한 학습환경에서 비교하여 ViT를 사용한 실시간 환경에서의 제스처 인식을 구현하고 실사용 가능성을 평가한다.
다국어 초록 (Multilingual Abstract)
This study explores the feasibility of a Vision Transformer (ViT)-based hand gesture recognition system for intuitive content manipulation in next-generation displays utilizing Holographic Optical Elements (HOEs). HOEs are optical components designed ...
This study explores the feasibility of a Vision Transformer (ViT)-based hand gesture recognition system for intuitive content manipulation in next-generation displays utilizing Holographic Optical Elements (HOEs). HOEs are optical components designed to alter the path of light at specific wavelengths. When integrated into immersive displays using HOE films, traditional input methods such as physical buttons or voice commands are often inadequate for delivering natural and intuitive user experiences. Instead, touchless 4D interaction systems are more suitable for enabling seamless interaction between users and holographic content. To address this, we propose a ViT-based hand gesture recognition system. Unlike conventional convolutional neural networks (CNNs), ViTs adopt a transformer architecture—originally developed for natural language processing—and apply global self-attention mechanisms to learn relationships across the entire input image without using local filters. In this work, we collect a custom RGB-D dataset using the Intel RealSense D455 sensor and compare the performance of ViT and CNN models under identical training conditions. We then evaluate the real-time recognition performance and practical applicability of the ViT-based system for immersive holographic interfaces.
차량 운전 환경에서 활용되는 음성 인터페이스 디자인 고려사항
중국과 소련 유화의 ‘민족화’ 경로에 대한 비교 연구: 20세기 중반을 중심으로