음성 상호작용은 LLM과 Audio Augmentation 도구의 확산으로 인해 더 많이 사용될 것이다 최근 대규모 언어모델의 발전은 사용자의 발화를 정교하게 해석하고 일관된 맥락으로 응답을 생성하는 능...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17450364
서울 : 서울대학교 대학원, 2026
학위논문(박사) -- 서울대학교 대학원 , 지능정보융합학과 지능정보융합전공 , 2026. 2
2026
영어
006.3
서울
XV, 196 ; 26 cm
지도교수: 이중식
I804:11032-000000193254
0
상세조회0
다운로드음성 상호작용은 LLM과 Audio Augmentation 도구의 확산으로 인해 더 많이 사용될 것이다 최근 대규모 언어모델의 발전은 사용자의 발화를 정교하게 해석하고 일관된 맥락으로 응답을 생성하는 능...
음성 상호작용은 LLM과 Audio Augmentation 도구의 확산으로 인해 더 많이 사용될 것이다 최근 대규모 언어모델의 발전은 사용자의 발화를 정교하게 해석하고 일관된 맥락으로 응답을 생성하는 능력을 끌어올렸고, 이에 더해 Audio Augmentation 도구의 발전과 함께 카메라·조도·모션 등 다양한 센서가 통합되면서 음성 인터랙션이 확장되고 있다. Apple HomePod, Google Home 등 차세대 기기들은 이러한 기술적 토대를 바탕으로 대화 능력을 고도화하고 있으며, 이는 더 이상 “듣고 말하는” 단일 모달리티가 아니라 사용자의 물리적·인지적 상태를 함께 해석하는 맥락지능형 시스템으로의 전환을 예고한다.
대화 품질 향상을 위해서는 Dull System이 아니라, Adaptive한 시스템 설계가 중요하다. 특히 사용자 Engagement를 중심에 둔 설계가 필수적이다. Engagement는 사용자가 상호작용 과정에서 인지적, 정서적, 행동적으로 얼마나 관여하는지의 정도를 가리키며, 흥미(interest), 몰입(immersion), 피드백(feedback), 새로움(novelty), 목표(goal), 정서(affect), 상호작용성(interactivity), 통제감(control), 내적 동기 (motivation), 기대(expectation) 등 복합 요인으로 구성된다. 따라서 음성 인터페이스에서 시스템이 사용자의 기대와 주의 상태를 제대로 파악해 적정한 양과 깊이의 응답을 Adaptive하게 제공하는 것은 대화 품질을 유의미하게 향상시킨다.
사용자 행동 중 Gaze를 센싱하여 Adaptive System을 설계해야 한다. 대화형 에이전트 맥락에서 사용자의 행동은 Engagement 추론을 위한 Proxy로 기능한다. 디스플레이가 제한적이거나 입력 채널이 협소한 Voice User Interface(VUI)에서는 발화 내용만으로 사용자의 참여 상태를 포착하기 어렵다. 이때 비언어적 행동 분석은 Engagement의 변화를 정교하게 추적할 수 있는 대안이며, 그 결과를 응답·피드백 전략에 반영하면 적응적이고 효과적인 상호작용을 구현할 수 있다. 본 논문은 이 중에서도 특히 Gaze 행동에 주목하여, Gaze가 대화 중 Engagement Transition 순간의 실시간 단서가 될 수 있는지를 실험적으로 검증한다.
음성상호작용에서 사용자 행동을 센싱할 때 Visual Cue를 제공하는 것은 중요하다. 본래 음성상호작용은 병행 인터액션으로 Auditory는 배경적이다. 운전, 요리, 걷기 등의 시나리오에서 소리는 배경적이며 Gaze와 같이 음성 대화중 발생하는 사용자 행동은 필연적으로 발생한다. 이를 통해 디바이스의 상태를 확인하거나, 사용자의 상태를 유추할 수 있다. 따라서 행동을 센싱하여 Adaptive System을 설계할 때 Visual Cue는 사용자가 기계에 주의를 주는 보조적 모드로 적극 사용돼야 한다.
본 연구의 목표는 사용자의 Gaze를 센싱해 Engagement Transition을 포착하고, 이를 응답 깊이의 전략과 시각적 큐를 조절하는 전략에 활용하는 것이 스마트스피커의 대화 품질을 향상시키는지 탐구하고 검증하는 것이다. 이를 검증하기 위해 세 번의 실험 연구를 수행하였다.
첫 번째 연구는 사용자의 Gaze 행동이 Engagement Transition을 어떻게 반영하는지를 이해하고, 이를 기반으로 스마트스피커의 응답에 대한 적응형 디자인 전략을 조사한다. 총 23명이 스트레스 상담형 대화에 참여했고, 339개 발화, 678개 Gaze 데이터, 그리고 113개의 인터뷰 피드백을 수집하였다. 실험은 발화 Gaze 조건을 Non-Gaze(NG), Once-Gaze(OG), Full-Gaze(FG)로 분류하여 발화 길이·단어 수·어휘 다양성·발화 연결 표현·지연 신호 등 발화 특성과의 관계를 정량 분석하고, 연속 발화 간 Gaze 조건 전환 패턴을 평가했다. 또한 사후 인터뷰 피드백을 3개의 축(User Satisfaction, Interaction Breakdown, Response Strategy)으로 코딩하였다. 분석 결과, 사용자의 Engagement 수준은 Gaze 행동이 변화하는, OG 조건에서 가장 높게 나타났다. 또한 NG와 FG에서는 Gaze 조건이 유지되는 패턴을 보인 반면, OG에서는 가장 역동적인 Gaze 조건 전환 패턴이 나타났다. 이는 Engagement Transition이 OG 조건의 발화에서 나타남을 시사한다. 사후 인터뷰에서도 사용자들은 스마트스피커를 바라보며 말할 때 더 깊고 구체적인 응답을 기대한 반면, 바라보지 않을 때는 간단하고 명료한 응답을 선호한다고 표현했다. 이러한 결과는 Gaze가 Engagement Transition의 가시적 지표일 뿐 아니라, 응답 깊이를 조절하는 전략적 단서로 활용 가능함이 시사한다.
두 번째 연구는 Gaze 행동 기반으로 Response 깊이를 조절하는 스마트스피커 Gaze-driven Adaptive Voice Interaction System (GAVIS) 를 개발하고, Engagement Transition을 고려한 실현 가능한 전략인지 검증하였다. 24명의 참여자가 여행 계획과 관련하여 정보탐색 대화를 수행하였고, 534개 발화와 1,068개 Gaze 데이터, 그리고 94개의 인터뷰 피드백을 수집하였다. 응답 조건은 시스템이 Gaze 행동과 상관없이 고정된 응답 깊이를 제공하는 두 개의 Control(Concise, in-Depth) 조건과 Gaze 행동에 따라 응답의 깊이를 조절하는 Switching 조건으로 설계하였다. 정량적 분석은 발화 길이, 단어 수, 어휘 다양성, 발화 연결 표현, 지연 신호를 비교했고, 연속 발화의 Gaze 전환 비율과 발화 전후 Gaze 변화율을 함께 살폈다. 정성적 분석은 사후 인터뷰 피드백을 4개 축(Information Appropriateness, Information Relevance, Interaction Alignment, User Engagement)으로 코딩해 사용자 인식과 기대를 도출했다. 분석 결과, Switching, in-Depth 조건 그룹은 Concise 조건 그룹보다 더 적은 턴 수의 대화를 하고 더 긴 발화가 나타났으며, 이는 Response 깊이를 조절하는 방식이 대화 품질을 변화시키는 전략임을 시사한다. 또한 모든 응답 조건에서 발화 종료 시점에 Gaze 비율이 증가하였으며, 특히 Switching 조건에서 가장 크게 증가하였다. 이는 Engagement Transition 순간 GAVIS와 교감하려는 Gaze 행동이 Switching 조건에서 가장 두드러졌음을 보여준다. 또한 사용자들은 Switching 조건의 응답 방식에 대해 전반적으로 긍정적으로 평가했다. 단, 상황에 따라 고정된 응답 방식을 선택할 수 있는 커스터마이징 옵션도 동시에 제공되기를 선호한다. 이러한 결과는 Gaze 행동에 따라 응답의 깊이를 조절하는 상호작용 방식은 실현 가능하며, Engagement Transition에 따라 전략적으로 작동하는 Adaptive한 시스템 설계가 가능함을 시사한다.
세 번째 연구는 스마트스피커가 Gaze 센싱에 대한 Visual Cue를 언제, 어떻게 제공하는 것이 적절한지 분석하고, Visual Cue에 대한 적응형 디자인 전략을 조사하였다. 12명의 참여자가 within-subjects 설계로 세 조건—Control(No Visual Cue), During-Utterance, During-Response—을 모두 경험했고, 575개 발화와 1,150개 Gaze 데이터, 그리고 184개 인터뷰 피드백을 수집하였다. 분석 결과, 사용자들은 Control 조건보다 Condition1, 2 조건에서 더 많은 Gaze 기반 발화와 Gaze 행동을 보였으며, Condition1, 2 조건이 Control 조건보다 더 반응적이고 상호작용을 느낀다고 표현했다. 또한 During-Utterance, During-Response 조건은 Control 조건보다 대화의 턴 수가 적게 나타났으며, 사용자들은 Visual Cue에 대하여 인지적 부담감을 표현했다. 또한 사용자들은 Visual Cue가 Response 중심으로 제공되는 During-Response 조건을 선호했다. 단, Response 전체가 아닌 발화 중 포착된 Gaze와 연결된 Semantic Moment에서만 Visual Cue가 제공되기를 선호한다고 표현했다. 이러한 결과는 Gaze 센싱 데이터를 기반으로 Visual Cue를 제공할 때는 적절한 전략이 필요함을 시사하며, 피드백 제공 타이밍의 적시성을 고려하여 시스템에 대한 사용자 신뢰와 상호작용 부담 간 균형을 이루는 설계 전략이 필요함을 시사한다.
세 번의 연구는 Gaze가 Engagement Transition을 포착하는 신뢰할 만한 proxy 지표이자, 응답과 시각적 큐를 전략적으로 조절할 수 있는 변수로서 유효함을 보여준다. 첫 번째 연구는 발화 경계에서의 OG/FG/NG 차이와 연속된 발화에서의 조건 전환 패턴을 통해 사용자의 Engagement Transition이 어떠한 Gaze 조건에서 발화를 할 때 나타나는지 근거를 제시했다. 두 번째 연구는 첫 번째 연구의 근거 위에 응답 깊이(Concise ↔ In-Depth)를 실시간으로 조절하는 Adaptive 시스템의 타당성을 설계하고 검증했다. 세 번째 연구는 Gaze 행동을 포착하고 시스템이 응답 전략을 바꾸는 과정을 시각화하여 투명성과 설명가능한 인터페이스에 대하여 논의하고 사용자 행동에 대한 시각적 큐의 타이밍과 맥락화 전략을 조사했다. 이 일련의 결과들은 오늘날 다양한 맥락에서 발생하는 음성 인터랙션이 사용자 행동과 맥락 정보에 적응적으로 반응하고 작동 방식을 설명가능한 방식으로 나타낼 때 대화 품질이 향상될 수 있음을 시사한다. 더불어 차량 인포테인먼트, 회의 보조, 홈 IoT 관리 등 Gaze 행동과 음성 정보를 동시에 고려해야 하는 다양한 환경에서의 확장 가능성도 논의된다.
본 논문은 Empirical, Strategical, Theoretical 측면에서 의미 있는 기여를 제공한다. 경험적으로, 세 단계 실험을 통해 Gaze 행동이 Engagement Transition을 반영하는 행동 지표임을 검증하고, 이를 응답 깊이를 조절하는 방식과 시각적 큐 제공 전략으로 연결해 실제 대화에서 작동함을 보였다. 전략적으로, 사용자의 Gaze 패턴을 단서로 응답을 간결 또는 구체적으로 전환하는 메커니즘과, 인지 부담을 최소화하면서 반응성을 전달하는 피드백 타이밍과 맥락화 전략을 조사했다. 또한 “Gaze 센싱 → Engagement Transition 포착 → 응답 적응 → 피드백 제공”으로 이어지는 설계 프레임을 제시해 행동 기반 이론과 인터랙션 디자인을 가교했고, 사용자 상태 기반 멀티모달, 적응형, 설명가능 인터랙션의 토대를 다졌다. 요약하면, 본 연구는 Gaze 행동을 참여 변화의 실시간 단서이자 설계 레버로 재정의함으로써, 스마트스피커의 대화 품질을 실질적으로 끌어올리는 인간 중심의 Gaze-driven Adaptive Voice Interaction을 학문적으로 정립하고 실무적으로 구체화했다.
다국어 초록 (Multilingual Abstract)
Voice interaction is expected to become increasingly prevalent, driven by the proliferation of Large Language Models (LLMs) and audio augmentation tools. Recent advancements in LLMs have significantly elevated the capability to precisely interpret use...
Voice interaction is expected to become increasingly prevalent, driven by the proliferation of Large Language Models (LLMs) and audio augmentation tools. Recent advancements in LLMs have significantly elevated the capability to precisely interpret user utterances and generate contextually consistent responses. Furthermore, the scope of voice interaction is expanding through the integration of diverse sensors—including cameras, ambient light, and motion sensors—alongside developments in audio augmentation. Next-generation devices, such as the Apple HomePod and Google Home, are leveraging these technological foundations to refine their conversational capabilities. This evolution signals a transition from a unimodal “listen-and-speak” paradigm to context-aware intelligent systems that simultaneously interpret the user’s physical and cognitive states.
To enhance interaction quality, it is critical to transition from static systems to adaptive system designs. In particular, a design framework centered on user engagement is essential. Engagement refers to the degree of a user's cognitive, emotional, and behavioral involvement during an interaction. It is a multidimensional construct composed of complex factors such as interest, immersion, feedback, novelty, goals, affect, interactivity, control, intrinsic motivation, and expectation. Therefore, in voice interfaces, the system's ability to accurately grasp the user's expectations and attentional state—and thereby provide adaptive responses of appropriate quantity and depth—significantly improves the quality of the dialogue.
Adaptive systems should be designed by sensing gaze behavior. In the context of conversational agents, user behavior serves as a proxy for inferring engagement. In Voice User Interfaces (VUIs), where displays are limited or input channels are restricted, capturing user engagement solely through verbal content is challenging. In this regard, analyzing non-verbal behavior offers a sophisticated alternative for precisely tracking engagement transitions. By incorporating these insights into response and feedback strategies, it is possible to realize adaptive and effective interaction. This study specifically focuses on gaze behavior to experimentally verify whether gaze can serve as a real-time cue for engagement transitions during conversation.
Providing visual cues is crucial when sensing user behavior in voice interaction. Inherently, voice interaction functions as a parallel interaction where the auditory modality often serves a background role. In scenarios such as driving, cooking, or walking, audio functions in the background, while user behaviors—such as gaze—inevitably occur during voice conversations. These behaviors facilitate the verification of device status or the inference of the user's state. Therefore, when designing adaptive systems based on behavioral sensing, visual cues must be actively utilized as an auxiliary mode to orient the user's attention toward the device.
The primary objective of this study is to investigate and verify whether sensing user gaze to capture engagement transitions—and subsequently utilizing this information for adaptive response depth and visual cue strategies—enhances the dialogue quality of smart speakers. To validate this hypothesis, three experimental studies were conducted.
The first study investigates how user gaze behaviors reflect engagement transitions and, based on these insights, explores adaptive design strategies for smart speaker responses. A total of 23 participants engaged in stress counseling conversations, yielding a dataset of 339 utterances, 678 gaze data points, and 113 interview feedbacks. The experiment classified gaze conditions during utterances into Non-Gaze (NG), Once-Gaze (OG), and Full-Gaze (FG) to quantitatively analyze their relationship with linguistic features such as utterance length, word count, lexical diversity, connectives, and delay signals. Furthermore, gaze condition transition patterns between consecutive utterances were evaluated. Additionally, post-interview feedback was coded along three axes: User Satisfaction, Interaction Breakdown, and Response Strategy. The analysis revealed that user engagement levels were highest in the OG condition, where gaze behavior shifts. While the NG and FG conditions showed patterns of maintaining the gaze condition, the OG condition exhibited the most dynamic gaze condition transition patterns. This implies that engagement transitions manifest in utterances under the OG condition. In post-experiment interviews, participants indicated that they expected deeper and more specific responses when addressing the smart speaker with gaze, whereas they preferred simple and concise responses when averting their gaze. These findings suggest that gaze serves not only as a visible indicator of engagement transitions but also as a strategic cue for adapting response depth.
The second study involved the development of the Gaze-driven Adaptive Voice Interaction System (GAVIS), which adapts response depth based on gaze behavior, and verified the feasibility of this strategy in consideration of engagement transitions. Twenty-four participants performed information-seeking conversations related to travel planning, yielding a dataset of 534 utterances, 1,068 gaze data points, and 94 interview feedbacks. The response conditions were designed with two Control conditions (Concise, In-depth), where the system provided a fixed response depth regardless of gaze behavior, and a Switching condition, where the system adapted response depth according to gaze behavior. Quantitative analysis compared utterance length, word count, lexical diversity, connectives, and delay signals, while also examining gaze transition rates between consecutive utterances and gaze change rates before and after utterances. Qualitative analysis coded post-interview feedback along four axes—Information Appropriateness, Information Relevance, Interaction Alignment, and User Engagement—to derive user perceptions and expectations. The analysis revealed that the Switching and In-depth condition groups engaged in conversations with fewer turns but longer utterances compared to the Concise condition group; this suggests that adapting response depth is a strategy that alters dialogue quality. Furthermore, the gaze ratio at the end of utterances increased across all response conditions, with the most significant increase observed in the Switching condition. This demonstrates that gaze behavior intended to interact with GAVIS during the moment of engagement transition was most prominent in the Switching condition. Additionally, users generally evaluated the response method of the Switching condition positively. However, they also expressed a preference for providing a simultaneous customization option to select a fixed response method depending on the situation. These findings suggest that an interaction method that adapts response depth based on gaze behavior is feasible, and that it is possible to design an adaptive system that operates strategically according to engagement transitions.
The third study analyzed when and how it is appropriate for a smart speaker to provide visual cues regarding gaze sensing, and investigated adaptive design strategies for these cues. Twelve participants experienced three conditions—Control (No Visual Cue), During-Utterance, and During-Response—in a within-subjects design, yielding a dataset of 575 utterances, 1,150 gaze data points, and 184 interview feedbacks. The analysis revealed that users exhibited more gaze-based utterances and gaze behaviors in the During-Utterance and During-Response conditions compared to the Control condition, and they described these conditions as feeling more responsive and interactive. Additionally, the During-Utterance and During-Response conditions resulted in fewer conversational turns than the Control condition, while users expressed cognitive burden regarding the visual cues. Users indicated a preference for the During-Response condition, where visual cues were provided primarily during the response phase. However, they specifically preferred visual cues to be displayed not throughout the entire response, but only at semantic moments linked to the gaze captured during the utterance. These findings suggest that appropriate strategies are requisite when providing visual cues based on gaze sensing data. Furthermore, they imply the necessity of design strategies that achieve a balance between user trust in the system and interaction burden, taking into account the timeliness of feedback provision.
Collectively, these three studies demonstrate that gaze serves as a reliable proxy for capturing engagement transitions and as a valid variable for strategically adapting responses and visual cues. The first study provided empirical evidence regarding the specific gaze conditions under which engagement transitions manifest during utterances, by analyzing differences among OG, FG, and NG conditions at utterance boundaries and examining transition patterns across consecutive utterances. Building upon the findings of the first study, the second study designed and verified the feasibility of an adaptive system that adapts response depth (Concise ↔ In-Depth) in real-time. The third study discussed transparency and explainable interfaces by visualizing the process of detecting gaze behaviors and shifting response strategies, while also investigating the timeliness and contextualization strategies for providing visual cues regarding user behavior. These collective findings suggest that the quality of dialogue in voice interactions, which occur in diverse contexts today, can be enhanced when systems respond adaptively to user behavior and contextual information, and when they present their operational mechanisms in an explainable manner. Furthermore, the study discusses the potential for extensibility in various environments where gaze behavior and voice information must be considered simultaneously, such as in-vehicle infotainment, meeting assistants, and home IoT management.
This dissertation offers significant contributions from empirical, strategic, and theoretical perspectives. Empirically, through three phases of experimental research, this study verified that gaze behavior serves as a behavioral indicator reflecting engagement transitions and demonstrated its viability in actual conversations by linking it to methods for adapting response depth and strategies for providing visual cues. Strategically, the study investigated mechanisms for switching responses between concise and in-depth formats using user gaze patterns as cues, as well as strategies for feedback timeliness and contextualization that convey responsiveness while minimizing cognitive burden. Furthermore, it presented a design framework proceeding from "gaze sensing" to "engagement transition capture," "response adaptation," and "feedback provision," thereby bridging behavioral theory with interaction design and laying the foundation for user state-based multimodal, adaptive, and explainable interactions. In summary, by redefining gaze behavior as both a real-time cue for engagement changes and a design lever, this study academically establishes and practically concretizes human-centered Gaze-driven Adaptive Voice Interaction, which substantially enhances the dialogue quality of smart speakers.
목차 (Table of Contents)