RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Building Multisensory Egocentric Intelligence: From Comprehensive to Proactive Perception = 다중감각 자기중심 지능 구축: 포괄적 지각에서 능동적 지각으로

    한글로보기

    https://www.riss.kr/link?id=T17449899

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    The long-standing human endeavor to augment perception has evolved from correcting sensory deficits to enhancing the capabilities of the general user population. However, existing intelligent systems often remain reactive, user-agnostic, and constrained by third-person, pre-recorded content, which limits their seamless integration into daily life. This thesis addresses these gaps by proposing Multisensory Egocentric Intelligence, a framework designed to comprehend multisensory context from a user-centric perspective and provide proactive assistance. This work establishes the core technologies for systems that should effectively interact with a user's perceptual input, comprehensively understand their internal context and surroundings, and proactively anticipate their future needs.
    This thesis focuses on three core components for Multisensory Egocentric Intelligence. Part I focuses on the interface for interaction, establishing a new benchmark for grounded audio-visual question answering in 360-degree panoramic videos to enable interactive spatial reasoning, and developing an unsupervised framework for omnidirectional saliency prediction for non-intrusive capture of user attention. Part II addresses comprehensive perception, assessing sensory plasticity by demonstrating 3D scene reconstruction from only binaural audio for the first time via a novel knowledge distillation technique, and proposing a novel framework to integrate heterogeneous modalities by stabilizing egocentric signals against user self-motion. Part III targets proactive assistance, introducing a model that forecasts a user's future visual attention in 3D to anticipate intent and presenting an efficient system that synthesizes appropriate spatial audio from visual context.
    Collectively, these contributions lay the groundwork for a new generation of assistive technologies and immersive extended reality experiences that perceive and interact with the world in a more human-like manner. The principles and models developed herein pave the way for future research into interactions in multi-party long-term conversation, improved reasoning capability with large multimodal models, and development of a unified perceptual world model.
    번역하기

    The long-standing human endeavor to augment perception has evolved from correcting sensory deficits to enhancing the capabilities of the general user population. However, existing intelligent systems often remain reactive, user-agnostic, and constrain...

    The long-standing human endeavor to augment perception has evolved from correcting sensory deficits to enhancing the capabilities of the general user population. However, existing intelligent systems often remain reactive, user-agnostic, and constrained by third-person, pre-recorded content, which limits their seamless integration into daily life. This thesis addresses these gaps by proposing Multisensory Egocentric Intelligence, a framework designed to comprehend multisensory context from a user-centric perspective and provide proactive assistance. This work establishes the core technologies for systems that should effectively interact with a user's perceptual input, comprehensively understand their internal context and surroundings, and proactively anticipate their future needs.
    This thesis focuses on three core components for Multisensory Egocentric Intelligence. Part I focuses on the interface for interaction, establishing a new benchmark for grounded audio-visual question answering in 360-degree panoramic videos to enable interactive spatial reasoning, and developing an unsupervised framework for omnidirectional saliency prediction for non-intrusive capture of user attention. Part II addresses comprehensive perception, assessing sensory plasticity by demonstrating 3D scene reconstruction from only binaural audio for the first time via a novel knowledge distillation technique, and proposing a novel framework to integrate heterogeneous modalities by stabilizing egocentric signals against user self-motion. Part III targets proactive assistance, introducing a model that forecasts a user's future visual attention in 3D to anticipate intent and presenting an efficient system that synthesizes appropriate spatial audio from visual context.
    Collectively, these contributions lay the groundwork for a new generation of assistive technologies and immersive extended reality experiences that perceive and interact with the world in a more human-like manner. The principles and models developed herein pave the way for future research into interactions in multi-party long-term conversation, improved reasoning capability with large multimodal models, and development of a unified perceptual world model.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    지각을 증강하려는 인류의 오랜 노력은 감각 결함을 교정하는 것에서부터 일반 사용자 집단의 능력을 향상시키는 것으로 발전해 왔다. 그러나 기존의 지능형 시스템은 종종 수동적이고, 사용자를 고려하지 않으며, 3인칭의 사전 녹화된 콘텐츠에 제약되어 일상 생활로의 원활한 통합에 한계를 보인다. 본 논문은 이러한 한계점들을 해결하기 위해 사용자 중심의 관점에서 다중감각적 맥락을 이해하고 능동적인 지원을 제공하도록 설계된 프레임워크인 다중감각 자기중심 지능 (Multisensory Egocentric Intelligence)을 제안한다. 본 연구는 사용자의 지각 입력과 효과적으로 상호작용하고, 사용자의 내적 맥락과 주변 환경을 포괄적으로 이해하며, 미래의 요구를 능동적으로 예측해야 하는 시스템을 위한 핵심 기술들을 정립한다.
    본 논문은 다중감각 자기중심 지능을 위한 세 가지 핵심 구성 요소에 초점을 맞춘다. 1부는 상호작용을 위한 인터페이스에 중점을 두며, 상호작용적 공간 추론을 가능하게 하기 위해 360도 파노라마 비디오에서의 근거 기반 시청각적 질의응답을 위한 새로운 벤치마크를 구축하고, 사용자의 주의를 비침입적으로 포착하기 위해 전방향 비디오 현저성 예측을 위한 비지도 프레임워크를 개발한다. 2부는 포괄적인 지각을 다루며, 새로운 지식 증류 기법을 통해 바이노럴 오디오만으로 3D 장면 재구성이 가능함을 최초로 입증하여 감각 가소성을 평가하고, 사용자의 자기 운동에 대해 자기중심적 신호를 안정화함으로써 이종의 신호 입력을 통합하는 새로운 프레임워크를 제안한다. 3부는 능동적인 지원을 목표로 하며, 의도를 예측하기 위해 사용자의 미래 시각적 주의를 3D로 예측하는 모델을 소개하고, 시각적 맥락으로부터 적절한 공간 오디오를 합성하는 효율적인 시스템을 제시한다.
    종합적으로, 이러한 기여들은 세상을 보다 인간과 유사한 방식으로 지각하고 상호작용하는 차세대 보조 기술 및 몰입형 확장 현실 경험의 토대를 마련한다. 본 연구에서 개발된 원리와 모델은 다자간 장기 대화에서의 상호작용, 대형 멀티모달 모델을 이용한 추론 능력 향상, 그리고 통합된 지각적 세계 모델 개발에 대한 향후 연구의 길을 열어준다.
    번역하기

    지각을 증강하려는 인류의 오랜 노력은 감각 결함을 교정하는 것에서부터 일반 사용자 집단의 능력을 향상시키는 것으로 발전해 왔다. 그러나 기존의 지능형 시스템은 종종 수동적이고, 사...

    지각을 증강하려는 인류의 오랜 노력은 감각 결함을 교정하는 것에서부터 일반 사용자 집단의 능력을 향상시키는 것으로 발전해 왔다. 그러나 기존의 지능형 시스템은 종종 수동적이고, 사용자를 고려하지 않으며, 3인칭의 사전 녹화된 콘텐츠에 제약되어 일상 생활로의 원활한 통합에 한계를 보인다. 본 논문은 이러한 한계점들을 해결하기 위해 사용자 중심의 관점에서 다중감각적 맥락을 이해하고 능동적인 지원을 제공하도록 설계된 프레임워크인 다중감각 자기중심 지능 (Multisensory Egocentric Intelligence)을 제안한다. 본 연구는 사용자의 지각 입력과 효과적으로 상호작용하고, 사용자의 내적 맥락과 주변 환경을 포괄적으로 이해하며, 미래의 요구를 능동적으로 예측해야 하는 시스템을 위한 핵심 기술들을 정립한다.
    본 논문은 다중감각 자기중심 지능을 위한 세 가지 핵심 구성 요소에 초점을 맞춘다. 1부는 상호작용을 위한 인터페이스에 중점을 두며, 상호작용적 공간 추론을 가능하게 하기 위해 360도 파노라마 비디오에서의 근거 기반 시청각적 질의응답을 위한 새로운 벤치마크를 구축하고, 사용자의 주의를 비침입적으로 포착하기 위해 전방향 비디오 현저성 예측을 위한 비지도 프레임워크를 개발한다. 2부는 포괄적인 지각을 다루며, 새로운 지식 증류 기법을 통해 바이노럴 오디오만으로 3D 장면 재구성이 가능함을 최초로 입증하여 감각 가소성을 평가하고, 사용자의 자기 운동에 대해 자기중심적 신호를 안정화함으로써 이종의 신호 입력을 통합하는 새로운 프레임워크를 제안한다. 3부는 능동적인 지원을 목표로 하며, 의도를 예측하기 위해 사용자의 미래 시각적 주의를 3D로 예측하는 모델을 소개하고, 시각적 맥락으로부터 적절한 공간 오디오를 합성하는 효율적인 시스템을 제시한다.
    종합적으로, 이러한 기여들은 세상을 보다 인간과 유사한 방식으로 지각하고 상호작용하는 차세대 보조 기술 및 몰입형 확장 현실 경험의 토대를 마련한다. 본 연구에서 개발된 원리와 모델은 다자간 장기 대화에서의 상호작용, 대형 멀티모달 모델을 이용한 추론 능력 향상, 그리고 통합된 지각적 세계 모델 개발에 대한 향후 연구의 길을 열어준다.

    더보기

    목차 (Table of Contents)

    • Contents
    • Abstract i
    • Chapter 1 Introduction 1
    • Chapter 2 Benchmarking Grounded Audio-Visual Question Answering 12
    • 2.1 Introduction 12
    • Contents
    • Abstract i
    • Chapter 1 Introduction 1
    • Chapter 2 Benchmarking Grounded Audio-Visual Question Answering 12
    • 2.1 Introduction 12
    • 2.2 Related Work 15
    • 2.3 Pano-AVQA Dataset 16
    • 2.3.1 Task Definition 17
    • 2.3.2 Data Collection 19
    • 2.3.3 Data Annotations 20
    • 2.3.4 Data Analysis 24
    • 2.4 Approach 26
    • 2.4.1 Input Representations 27
    • 2.4.2 Encoder 29
    • 2.4.3 Training 30
    • 2.5 Experiments 32
    • 2.5.1 Experimental Setup 32
    • 2.5.2 Results and Analyses 32
    • 2.6 Conclusion 37
    • Chapter 3 Omnidirectional Video Saliency Detection without Supervision 39
    • 3.1 Introduction 39
    • 3.2 Related Work 42
    • 3.3 Approach 45
    • 3.3.1 The Encoder for 360◦ Videos 46
    • 3.3.2 Spatiotemporal Fusion 49
    • 3.3.3 The Decoder for Saliency Map 50
    • 3.3.4 Learning Objectives 51
    • 3.4 Experiments 52
    • 3.4.1 Experiment Setting 53
    • 3.4.2 Results on Saliency Detection 55
    • 3.4.3 Omnidirectional Video Quality Assessment 62
    • 3.4.4 Qualitative Results63
    • 3.5 Conclusion 65
    • Chapter 4 2D-3D Scene Reconstruction from Binaural Inputs without Seeing 68
    • 4.1 Introduction 68
    • 4.2 Related Work 71
    • 4.3 Approach 73
    • 4.3.1 Vision-to-Audio Knowledge Distillation 74
    • 4.3.2 Spatial Alignment via Matching 75
    • 4.3.3 Training and Inference 78
    • 4.4 Experiments 80
    • 4.4.1 The DAPS Benchmark 81
    • 4.4.2 Results of Depth Estimation 82
    • 4.4.3 Results of Semantic Segmentation 87
    • 4.4.4 Results of 3D Scene Reconstruction 89
    • 4.5 Conclusion 93
    • Chapter 5 Self-Motion-Aware Multisensory Localization and Anticipation 95
    • 5.1 Introduction 95
    • 5.2 Related Work 98
    • 5.3 Spherical World-Locking 100
    • 5.3.1 Explicit Spherical World-Locking 101
    • 5.3.2 Implicit Spherical World-Locking 102
    • 5.4 Multisensory Spherical World-Locked Transformer 104
    • 5.4.1 MuST Encoder 104
    • 5.4.2 MuST Decoder 106
    • 5.5 Experiments 107
    • 5.5.1 Audio-Visual Active Speaker Localization 108
    • 5.5.2 Auditory Spherical Source Localization 111
    • 5.5.3 Egocentric Behavior Anticipation 112
    • 5.6 Discussion 115
    • 5.7 Conclusion 117
    • Chapter 6 Predictive Gaze Modeling in 3D Observations 119
    • 6.1 Introduction 119
    • 6.2 Related Work 122
    • 6.3 EgoSpanLift: Lifting Egocentric Gaze Prediction from 2D to 3D 124
    • 6.3.1 Preliminary: Simultaneous Localization and Mapping 124
    • 6.3.2 Keypoint Selection and Classification 125
    • 6.3.3 Multi-level Volumetric Region Localization 128
    • 6.4 Forecasting Network 129
    • 6.5 Benchmarking Egocentric 3D Visual Span Forecasting 133
    • 6.5.1 Forecasting Egocentric Daily Activities 133
    • 6.5.2 Forecasting Skilled Activites at Scale 137
    • 6.5.3 Extension to 2D Gaze Anticipation 139
    • 6.6 Discussion 141
    • 6.7 Conclusion 144
    • Chapter 7 Efficient Scene-Aware Video to Spatial Audio Generation 148
    • 7.1 Introduction 148
    • 7.2 Related Work 151
    • 7.2.1 Video-to-Audio Generation 151
    • 7.2.2 Audio Spatialization with Visual Cues 152
    • 7.2.3 Audio Generation with Neural Audio Codecs 153
    • 7.3 Video-to-Ambisonics Generation 153
    • 7.3.1 Background: First-Order Ambisonics 153
    • 7.3.2 Task Description 154
    • 7.3.3 Evaluation Metrics 155
    • 7.3.4 Dataset 157
    • 7.3.5 Filtering Strategy Analysis 160
    • 7.4 Approach 162
    • 7.4.1 Input Representation 162
    • 7.4.2 Autoregressive Generation 164
    • 7.4.3 Single Codebook-Based Autoregressive Generation 166
    • 7.4.4 Postprocessing and Inference-Time Refinement 168
    • 7.5 Experiments 170
    • 7.5.1 Evaluation Protocol 170
    • 7.5.2 Quantitative analysis 171
    • 7.5.3 Qualitative Analysis 175
    • 7.6 Conclusion 176
    • Chapter 8 Conclusion 180
    • 8.1 Summary of Contribution 180
    • 8.2 Future Work 183
    • Acknowledgements 230
    • 요약 232
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼