RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    생성형 AI의 자연어 추론을 통한 음성 인식 게임 상호작용 UX에 관한 연구 : 판타지 장르를 중심으로 = A Study on Game Interaction UX based on Speech Recognition using Generative AI's Natural Language Inference

    한글로보기

    https://www.riss.kr/link?id=T17400946

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This study addresses the phenomenon where, despite the rapid advancement of Generative AI and Large Language Model (LLM) technologies contributing to the efficiency of resource production in the game industry, innovation in the actual player interaction experience (UX) remains insufficient. In particular, existing voice recognition games rely on 'Keyword Spotting' methods that recognize only pre-defined specific commands. Consequently, they fail to accommodate players' free utterances and force the memorization of precise commands, thereby causing high cognitive load and hindering immersion. This study aims to overcome these structural limitations.
    Accordingly, this study proposes a 'Generative AI-based Intent Inference Interaction System' that utilizes the advanced context inference capabilities of Large Language Models to interpret players' imperfect natural language utterances as 'Intent' rather than simple errors and connect them to specific ingame actions, verifying its effectiveness empirically.
    To conduct this research, the rigid interaction problems of existing systems and the diverse utterance desires of players were first identified through literature review and Video Analysis. Based on this, a real-time testbed was established by integrating the Gemini API and Azure Speech API within an Unreal Engine 5 environment. Furthermore, to suppress AI Hallucination and output structured data processable by the game engine, a 'Goal-Oriented Constraint Prompt,' combining system goal definition and constraint setting, was designed through prompt engineering.
    The results of a Usability Test conducted with 10 adults to verify the efficacy of the proposed system are as follows. First, quantitative data analysis showed that the proposed system improved task performance efficiency by approximately 51.8% compared to the existing keyword method, and drastically reduced the average number of attempts per unit task from 2.43 (SD=1.58) to 1.17 (SD=0.38) (t=8.006, p < 0.001). In particular, the significant reduction in standard deviation implies that the proposed system provides consistent interaction performance regardless of individual user differences or speech habits. Second, qualitative interview and speech data analysis confirmed that players' utterance patterns were not chaotic noise but could be systematized into three major categories: Pleonasm Speech, Deconstruction Speech, and Distort Speech, along with seven detailed sub-types. Participants experienced 'immersive chanting' as if they were actual wizards, freed from the compulsion to input correct commands. They showed a positive attitude change, perceiving the interaction itself as playful content, such as utilizing memes or intentionally distorting pronunciation.
    In conclusion, this study demonstrates the feasibility of an 'Intent-Centric' interface that accepts players' natural language habits as they are. Based on this, the study presents a classification system for user utterance types, along with prompt engineering and UX design guidelines for applying Generative AI to games. The results of this study are expected to make academic and industrial contributions as a practical Reference Model for maximizing user experience beyond simple technical implementation in the future development of AI-based game content
    번역하기

    This study addresses the phenomenon where, despite the rapid advancement of Generative AI and Large Language Model (LLM) technologies contributing to the efficiency of resource production in the game industry, innovation in the actual player interacti...

    This study addresses the phenomenon where, despite the rapid advancement of Generative AI and Large Language Model (LLM) technologies contributing to the efficiency of resource production in the game industry, innovation in the actual player interaction experience (UX) remains insufficient. In particular, existing voice recognition games rely on 'Keyword Spotting' methods that recognize only pre-defined specific commands. Consequently, they fail to accommodate players' free utterances and force the memorization of precise commands, thereby causing high cognitive load and hindering immersion. This study aims to overcome these structural limitations.
    Accordingly, this study proposes a 'Generative AI-based Intent Inference Interaction System' that utilizes the advanced context inference capabilities of Large Language Models to interpret players' imperfect natural language utterances as 'Intent' rather than simple errors and connect them to specific ingame actions, verifying its effectiveness empirically.
    To conduct this research, the rigid interaction problems of existing systems and the diverse utterance desires of players were first identified through literature review and Video Analysis. Based on this, a real-time testbed was established by integrating the Gemini API and Azure Speech API within an Unreal Engine 5 environment. Furthermore, to suppress AI Hallucination and output structured data processable by the game engine, a 'Goal-Oriented Constraint Prompt,' combining system goal definition and constraint setting, was designed through prompt engineering.
    The results of a Usability Test conducted with 10 adults to verify the efficacy of the proposed system are as follows. First, quantitative data analysis showed that the proposed system improved task performance efficiency by approximately 51.8% compared to the existing keyword method, and drastically reduced the average number of attempts per unit task from 2.43 (SD=1.58) to 1.17 (SD=0.38) (t=8.006, p < 0.001). In particular, the significant reduction in standard deviation implies that the proposed system provides consistent interaction performance regardless of individual user differences or speech habits. Second, qualitative interview and speech data analysis confirmed that players' utterance patterns were not chaotic noise but could be systematized into three major categories: Pleonasm Speech, Deconstruction Speech, and Distort Speech, along with seven detailed sub-types. Participants experienced 'immersive chanting' as if they were actual wizards, freed from the compulsion to input correct commands. They showed a positive attitude change, perceiving the interaction itself as playful content, such as utilizing memes or intentionally distorting pronunciation.
    In conclusion, this study demonstrates the feasibility of an 'Intent-Centric' interface that accepts players' natural language habits as they are. Based on this, the study presents a classification system for user utterance types, along with prompt engineering and UX design guidelines for applying Generative AI to games. The results of this study are expected to make academic and industrial contributions as a practical Reference Model for maximizing user experience beyond simple technical implementation in the future development of AI-based game content

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    본 연구는 최근 생성형 AI(Generative AI)와 거대 언어 모델(LLM) 기술의 급격한 발전이 게임 산업의 리소스 생산 효율화에 기여하고 있음에도 불구하고, 실제 플레이어의 상호작용 경험(UX) 혁신에는 미진한 현상에 주목하였다. 특히 기존 음성 인식 게임이 사전에 정의된 특정 명령어만을 인식하는 ‘키워드 스포팅(Keyword Spotting)’ 방식에 의존함에 따라, 플레이어의 자유로운 발화를 수용하지 못하고 정확한 명령어 암기를 강요하여 높은 인지적 부하(Cognitive Load)와 몰입 저해를 유발한다는 구조적 한계를 극복하고자 하였다.
    이에 본 연구는 거대 언어 모델의 고도화된 문맥 추론 능력을 활용하여, 플레이어의 불완전한 자연어 발화를 단순한 오류가 아닌 ‘의도(Intent)’로 해석하고 이를 게임 내 구체적인 액션으로 연결하는 ‘생성형 AI 기반 의도 추론 상호작용 시스템’을 제안하고 그 효용성을 실증적으로 검증하였다.
    연구 수행을 위해 우선 문헌 연구와 비디오 분석(Video Analysis)을 통해 기존 시스템의 경직된 상호작용 문제와 플레이어의 다양한 발화 욕구를 규명하였다. 이를 바탕으로 Unreal Engine 5 환경에서 Gemini API 와 Azure Speech API 를 연동한 실시간 테스트 베드를 구축하였으며, AI 의 Hallucination 현상을 억제하고 게임 엔진이 처리 가능한 정형 데이터를 출력하기 위해 시스템 목표 정의와 제약 조건 설정이 결합된 ‘목적 지향적 제약 조건 프롬프트(Goal-Oriented Constraint Prompt)’를 프롬프트 엔지니어링하였다.
    제안 시스템의 효용성을 검증하기 위해 성인 10 명을 대상으로 사용자 평가 (Usability Test)를 수행한 결과는 다음과 같다. 첫째, 정량적 데이터 분석 결과 제안 시스템은 기존 키워드 방식 대비 과업 수행의 효율성을 약 51.8% 향상시켰으며, 단위 과업당 평균 시도 횟수를 2.43 회(SD=1.58)에서 1.17 회(SD=0.38)로 획기적으로 감소시켰다(t=8.006, p < 0.001). 특히 표준편차의 유의미한 감소는 제안 시스템이 사용자의 개인차나 발화 습관에 구애를 받지 않고 일관된 상호작용 성능을 제공함을 시사한다. 둘째, 정성적 인터뷰 및 발화 데이터 분석 결과, 플레이어의 발화 패턴은 무질서한 잡음이 아닌 Pleonasm Speech, Deconstruction Speech, Distort Speech 의 3대 범주와 그 하위의 7가지 세부 유형으로 체계화됨을 확인하였다.
    참여자들은 제안 시스템을 통해 정확한 명령어를 입력해야 한다는 강박에서 벗어나 실제 마법사가 된 듯한 ‘몰입적 영창’을 경험하였으며, 밈(Meme)을 활용하거나 발음을 의도적으로 변형하는 등 상호작용 자체를 유희적 콘텐츠로 인식하는 긍정적인 태도 변화를 보였다.
    결론적으로 본 연구는 플레이어의 자연어 습관을 있는 그대로 수용하는 ‘의도 중심(Intent-Centric)’ 인터페이스의 실현 가능성을 입증하였으며, 이를 바탕으로 생성형 AI 게임 적용을 위한 사용자 발화 유형 분류 체계와 프롬프트 엔지니어링 및 UX 디자인 가이드라인을 제시하였다. 본 연구의 결과는 향후 AI 를 활용한 게임 콘텐츠 개발에 있어, 단순한 기술적 구현을 넘어 사용자 경험을 극대화하기 위한 실질적인 참조 모델(Reference Model)로서 학술적 및 산업적 기여를 할 것으로 기대된다
    번역하기

    본 연구는 최근 생성형 AI(Generative AI)와 거대 언어 모델(LLM) 기술의 급격한 발전이 게임 산업의 리소스 생산 효율화에 기여하고 있음에도 불구하고, 실제 플레이어의 상호작용 경험(UX) 혁신에...

    본 연구는 최근 생성형 AI(Generative AI)와 거대 언어 모델(LLM) 기술의 급격한 발전이 게임 산업의 리소스 생산 효율화에 기여하고 있음에도 불구하고, 실제 플레이어의 상호작용 경험(UX) 혁신에는 미진한 현상에 주목하였다. 특히 기존 음성 인식 게임이 사전에 정의된 특정 명령어만을 인식하는 ‘키워드 스포팅(Keyword Spotting)’ 방식에 의존함에 따라, 플레이어의 자유로운 발화를 수용하지 못하고 정확한 명령어 암기를 강요하여 높은 인지적 부하(Cognitive Load)와 몰입 저해를 유발한다는 구조적 한계를 극복하고자 하였다.
    이에 본 연구는 거대 언어 모델의 고도화된 문맥 추론 능력을 활용하여, 플레이어의 불완전한 자연어 발화를 단순한 오류가 아닌 ‘의도(Intent)’로 해석하고 이를 게임 내 구체적인 액션으로 연결하는 ‘생성형 AI 기반 의도 추론 상호작용 시스템’을 제안하고 그 효용성을 실증적으로 검증하였다.
    연구 수행을 위해 우선 문헌 연구와 비디오 분석(Video Analysis)을 통해 기존 시스템의 경직된 상호작용 문제와 플레이어의 다양한 발화 욕구를 규명하였다. 이를 바탕으로 Unreal Engine 5 환경에서 Gemini API 와 Azure Speech API 를 연동한 실시간 테스트 베드를 구축하였으며, AI 의 Hallucination 현상을 억제하고 게임 엔진이 처리 가능한 정형 데이터를 출력하기 위해 시스템 목표 정의와 제약 조건 설정이 결합된 ‘목적 지향적 제약 조건 프롬프트(Goal-Oriented Constraint Prompt)’를 프롬프트 엔지니어링하였다.
    제안 시스템의 효용성을 검증하기 위해 성인 10 명을 대상으로 사용자 평가 (Usability Test)를 수행한 결과는 다음과 같다. 첫째, 정량적 데이터 분석 결과 제안 시스템은 기존 키워드 방식 대비 과업 수행의 효율성을 약 51.8% 향상시켰으며, 단위 과업당 평균 시도 횟수를 2.43 회(SD=1.58)에서 1.17 회(SD=0.38)로 획기적으로 감소시켰다(t=8.006, p < 0.001). 특히 표준편차의 유의미한 감소는 제안 시스템이 사용자의 개인차나 발화 습관에 구애를 받지 않고 일관된 상호작용 성능을 제공함을 시사한다. 둘째, 정성적 인터뷰 및 발화 데이터 분석 결과, 플레이어의 발화 패턴은 무질서한 잡음이 아닌 Pleonasm Speech, Deconstruction Speech, Distort Speech 의 3대 범주와 그 하위의 7가지 세부 유형으로 체계화됨을 확인하였다.
    참여자들은 제안 시스템을 통해 정확한 명령어를 입력해야 한다는 강박에서 벗어나 실제 마법사가 된 듯한 ‘몰입적 영창’을 경험하였으며, 밈(Meme)을 활용하거나 발음을 의도적으로 변형하는 등 상호작용 자체를 유희적 콘텐츠로 인식하는 긍정적인 태도 변화를 보였다.
    결론적으로 본 연구는 플레이어의 자연어 습관을 있는 그대로 수용하는 ‘의도 중심(Intent-Centric)’ 인터페이스의 실현 가능성을 입증하였으며, 이를 바탕으로 생성형 AI 게임 적용을 위한 사용자 발화 유형 분류 체계와 프롬프트 엔지니어링 및 UX 디자인 가이드라인을 제시하였다. 본 연구의 결과는 향후 AI 를 활용한 게임 콘텐츠 개발에 있어, 단순한 기술적 구현을 넘어 사용자 경험을 극대화하기 위한 실질적인 참조 모델(Reference Model)로서 학술적 및 산업적 기여를 할 것으로 기대된다

    더보기

    목차 (Table of Contents)

    • 1.서론 7
    • 1.1.연구 배경 7
    • 1.2.연구 목적 10
    • 1.3.연구방법 및 구성 12
    • 2.이론적 배경 16
    • 1.서론 7
    • 1.1.연구 배경 7
    • 1.2.연구 목적 10
    • 1.3.연구방법 및 구성 12
    • 2.이론적 배경 16
    • 2.1.생성형 AI와 거대 언어 모델(LLM) 16
    • 2.1.1.거대 언어 모델의 정의 및 기술적 특성 16
    • 2.1.2.LLM의 창발적 능력과 추론 능력 17
    • 2.2.게임 내 음성 인식 상호작용 19
    • 2.2.1.음성 인식(STT) 기술의 정의 및 현황 19
    • 2.2.2.기존 음성 게임의 상호작용 방식과 한계 20
    • 2.2.3.음성 인식을 활용한 게임의 사례 및 한계 21
    • 2.2.4.음성 인식과 AI의 결합: 문맥 기반 후처리 및 의도 추론 23
    • 3.생성형 AI 기반 상호작용 시스템 설계 26
    • 3.1.Trend Research 26
    • 3.1.1.Secondary Research & Competitors Analysis 26
    • 3.1.2.Domain Knowledge 파악 28
    • 3.1.3.Video Analysis 29
    • 3.1.4.현상 분석 31
    • 3.2.실험 설계 32
    • 3.2.1.프롬프트 제작 32
    • 3.2.2.Gemini를 통한 프롬프트 1차 검증 37
    • 3.3.시스템 아키텍처 및 테스트 빌드 구현 42
    • 3.3.1.시스템 아키텍처 및 Main Frame 42
    • 3.3.2.개발 환경 및 도구 선정 43
    • 3.3.3.언리얼 블루프린트(Blueprint) 기능을 통한 게임 구현 44
    • 4.실험 및 결과 분석 49
    • 4.1.실험 설계 및 절차 49
    • 4.1.1.실험 참가자 및 조건 49
    • 4.1.2.실험 과제(Task) 및 절차 50
    • 4.1.3.주문 리스트 휴리스틱 유형 51
    • 4.2.실험 결과 및 분석 51
    • 4.2.1.Pleonasm Speech 51
    • 4.2.2.Deconstruction Speech 55
    • 4.2.3.Distort Speech 58
    • 4.2.4.시스템별 주문 인식 정확도 및 과업 효율성 비교 61
    • 4.2.5.사용자 심층 인터뷰 분석 64
    • 4.3.결과 검증 및 논의 66
    • 4.3.1.가설의 검증: 효율성과 정확도의 상관관계 66
    • 4.3.2.‘암기'에서 '표현'으로의 사용자 경험 전이 67
    • 4.3.3.종합 논의 67
    • 5.생성형 AI 게임 상호작용 디자인 가이드라인 69
    • 5.1.7가지 사용자 발화 유형 69
    • 5.2.프롬프트 엔지니어링 및 기술적 가이드라인 72
    • 5.3.생성형 AI 게임 상호작용을 위한 UX 디자인 가이드라인 76
    • 6.결론 79
    • 6.1.연구 요약 79
    • 6.2.연구의 의의 및 기대효과 80
    • 6.3.한계점 및 향후 연구 과제 81
    • 참고문헌Reference 83
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼