Reliable and efficient navigation in complex and dynamic environments remains a core challenge in robotics.
Central to this challenge is the need for robust, compact, and spatially consistent scene representations that enable embodied agents to local...
Reliable and efficient navigation in complex and dynamic environments remains a core challenge in robotics.
Central to this challenge is the need for robust, compact, and spatially consistent scene representations that enable embodied agents to localize, explore, and adapt to diverse conditions.
This dissertation introduces a suite of scalable and adaptive mapping approaches that support long-term, vision-based navigation.
By addressing domain adaptation, multi-agent exploration, and cross-modal scene understanding, the proposed representations provide a unified foundation for building generalizable and context-aware navigation systems in visually and structurally diverse environments.
To this end, the dissertation proposes map-centric methods that span multiple spatial levels, from low-level grid maps to high-level graphs, and integrate data from diverse modalities.
These methods address key challenges such as domain generalization, structural consistency, multi-agent coordination, and long-term adaptability.
Together, they form a cohesive framework for building spatial representations that are robust, scalable, and transferable across domains, agents, and time.
To support robust and scalable navigation, this dissertation introduces three key contributions.
First, it proposes a self-supervised adaptation method that handles domain shifts by enforcing structural consistency in grid maps.
This method combines a consistency loss to stabilize mapping under noisy poses with a style-transfer refinement that denoises maps,
and uses a curriculum learning strategy to correct visual and pose noise through multi-scale spatial reasoning.
Second, it presents a decentralized multi-agent exploration framework, where agents coordinate via similarity score maps—spatial representations quantifying visual novelty relative to others’ observations.
By exchanging local topological graphs and identifying dissimilar frontiers, agents explore efficiently without pose sharing or global map fusion.
Finally, it introduces a hybrid region-based map representation that segments the environment into regions, each storing metric, visual, and 3D information.
This modular graph structure enables multi-scale reasoning, cross-session alignment, and persistent mapping across heterogeneous sensors, supporting flexible and generalizable navigation in dynamic environments.
Collectively, this dissertation presents a unified set of methods for constructing adaptive, robust, and scalable map representations to support vision-based navigation.
These contributions address core challenges in domain adaptation, multi-agent collaboration, and hybrid spatial reasoning, advancing the autonomy, generalization, and long-term reliability of embodied agents in complex, real-world settings.