The long-standing human endeavor to augment perception has evolved from correcting sensory deficits to enhancing the capabilities of the general user population. However, existing intelligent systems often remain reactive, user-agnostic, and constrain...
The long-standing human endeavor to augment perception has evolved from correcting sensory deficits to enhancing the capabilities of the general user population. However, existing intelligent systems often remain reactive, user-agnostic, and constrained by third-person, pre-recorded content, which limits their seamless integration into daily life. This thesis addresses these gaps by proposing Multisensory Egocentric Intelligence, a framework designed to comprehend multisensory context from a user-centric perspective and provide proactive assistance. This work establishes the core technologies for systems that should effectively interact with a user's perceptual input, comprehensively understand their internal context and surroundings, and proactively anticipate their future needs.
This thesis focuses on three core components for Multisensory Egocentric Intelligence. Part I focuses on the interface for interaction, establishing a new benchmark for grounded audio-visual question answering in 360-degree panoramic videos to enable interactive spatial reasoning, and developing an unsupervised framework for omnidirectional saliency prediction for non-intrusive capture of user attention. Part II addresses comprehensive perception, assessing sensory plasticity by demonstrating 3D scene reconstruction from only binaural audio for the first time via a novel knowledge distillation technique, and proposing a novel framework to integrate heterogeneous modalities by stabilizing egocentric signals against user self-motion. Part III targets proactive assistance, introducing a model that forecasts a user's future visual attention in 3D to anticipate intent and presenting an efficient system that synthesizes appropriate spatial audio from visual context.
Collectively, these contributions lay the groundwork for a new generation of assistive technologies and immersive extended reality experiences that perceive and interact with the world in a more human-like manner. The principles and models developed herein pave the way for future research into interactions in multi-party long-term conversation, improved reasoning capability with large multimodal models, and development of a unified perceptual world model.