RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Spatial and Linguistic Understanding for Autonomous Navigation in Smart Cities = 스마트시티 자율주행을 위한 공간 및 언어 이해

    한글로보기

    https://www.riss.kr/link?id=T17450951

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    스마트시티 환경에서의 자율주행은 불완전한 센싱, 부정확하거나 불완전한 지도 정보, 그리고 다양하고 비정형적인 사용자 명령이 공존하는 지속적인 환경적 제약 하에서도 강건한 동작을 요구하였다. 이러한 조건은 극심한 기상 변화, 제한된 가시성, 신호 감쇠, 그리고 센서 성능 저하로부터 발생하였으며, 자율 이동 시스템의 신뢰성을 근본적으로 저해하였다. 이 과정에서 공간 정보, 이종 모달 정보, 그리고 언어 정보를 포함한 이질적이고 불완전한 입력을 자율주행을 위한 과제 관련 표현으로 변환하는 것이 핵심적인 어려움으로 작용하였다.

    본 학위논문은 이러한 문제를 표현 기반의 통합적 관점에서 다루었으며, 자율주행을 지도 작성, 위치 추정, 그리고 경로 계획으로 구성된 순차적 파이프라인으로 정의하였다. 각 단계에서의 실패를 개별적으로 다루기보다, 기존 표현 방식이 환경적 제약 하에서 취약해지고 기하 정보의 열화, 이종 모달 간 불일치, 그리고 언어 정보의 비효율적 활용을 충분히 연결하지 못한다는 공통된 근본 문제를 식별하였다. 이를 해결하기 위해, 본 논문은 자율주행 파이프라인의 각 단계에 특화된 과제 지향적 표현을 구성하면서도 단계 간 개념적 일관성을 유지하는 통합적 프레임워크를 제안하였다.

    지도 작성 단계에서는 기하적 관측 가능성이 심각하게 제한된 환경을 대상으로 하였다. 희소하고 잡음이 많은 포인트 클라우드는 구조적 부트스트래핑을 통해 보강되었으며, 기하, 법선, 곡률 통계로부터 추출된 상호 보완적인 구조적 특징을 활용하여 회전에 불변한 구조적 표현을 형성하였다. 이러한 표현은 구조적 유사도에 기반한 신뢰성 있는 루프 검출을 가능하게 하였고, 고해상도 센서, 키포인트 추출, 또는 학습 기반 전처리에 의존하지 않으면서 전역적으로 일관된 지도를 구축할 수 있도록 하였다.

    이러한 표현 중심의 기반 위에서, 위치 추정 단계에서는 자연어 설명과 공간 지도 데이터 간의 표현 불일치 문제가 다루어졌다. 텍스트 입력과 OpenStreetMap 데이터는 통합된 씬 그래프 표현으로 정규화되었으며, 이를 통해 공통 임베딩 공간에서 이종 모달 간 비교가 가능해졌다. 이러한 정규화 과정은 지도 정보가 불완전하거나 최신 상태가 아니거나 원시 센서 관측과의 상관성이 약한 환경에서도 신뢰성 있는 텍스트 기반 장소 인식을 가능하게 하였다.

    경로 계획 단계에서는 언어적 의도를 표현 학습의 대상으로 확장하였다. 자연어 명령은 암묵적인 선호와 제약 조건을 추출하도록 해석되었으며, 이는 그래프 기반 경로 계획기의 가중치 함수에 직접적으로 인코딩되었다. 이러한 의도 인지 표현을 OpenStreetMap 기반 경로 탐색에 통합함으로써, 생성된 경로는 최단 경로 최적화를 넘어 사용자 의도, 안전성, 그리고 상황적 맥락을 보다 잘 반영하도록 하였다.

    종합적으로, 본 학위논문은 환경적 제약 하에서 발생하는 자율주행 문제를 지도 작성, 위치 추정, 그리고 경로 계획 전반에 걸친 과제 지향적 표현 구성의 문제로 체계화하였다. 제안된 프레임워크는 기하적, 이종 모달, 그리고 언어적 추론을 통합적으로 연결함으로써 인간 중심적이며 강건한 자율주행의 가능성을 제시하였으며, 불확실성과 다양성, 그리고 대규모 특성을 갖는 실제 스마트시티 환경에서의 신뢰성 있는 자율 이동 시스템 구현 가능성을 입증하였다.
    번역하기

    스마트시티 환경에서의 자율주행은 불완전한 센싱, 부정확하거나 불완전한 지도 정보, 그리고 다양하고 비정형적인 사용자 명령이 공존하는 지속적인 환경적 제약 하에서도 강건한 동작을 ...

    스마트시티 환경에서의 자율주행은 불완전한 센싱, 부정확하거나 불완전한 지도 정보, 그리고 다양하고 비정형적인 사용자 명령이 공존하는 지속적인 환경적 제약 하에서도 강건한 동작을 요구하였다. 이러한 조건은 극심한 기상 변화, 제한된 가시성, 신호 감쇠, 그리고 센서 성능 저하로부터 발생하였으며, 자율 이동 시스템의 신뢰성을 근본적으로 저해하였다. 이 과정에서 공간 정보, 이종 모달 정보, 그리고 언어 정보를 포함한 이질적이고 불완전한 입력을 자율주행을 위한 과제 관련 표현으로 변환하는 것이 핵심적인 어려움으로 작용하였다.

    본 학위논문은 이러한 문제를 표현 기반의 통합적 관점에서 다루었으며, 자율주행을 지도 작성, 위치 추정, 그리고 경로 계획으로 구성된 순차적 파이프라인으로 정의하였다. 각 단계에서의 실패를 개별적으로 다루기보다, 기존 표현 방식이 환경적 제약 하에서 취약해지고 기하 정보의 열화, 이종 모달 간 불일치, 그리고 언어 정보의 비효율적 활용을 충분히 연결하지 못한다는 공통된 근본 문제를 식별하였다. 이를 해결하기 위해, 본 논문은 자율주행 파이프라인의 각 단계에 특화된 과제 지향적 표현을 구성하면서도 단계 간 개념적 일관성을 유지하는 통합적 프레임워크를 제안하였다.

    지도 작성 단계에서는 기하적 관측 가능성이 심각하게 제한된 환경을 대상으로 하였다. 희소하고 잡음이 많은 포인트 클라우드는 구조적 부트스트래핑을 통해 보강되었으며, 기하, 법선, 곡률 통계로부터 추출된 상호 보완적인 구조적 특징을 활용하여 회전에 불변한 구조적 표현을 형성하였다. 이러한 표현은 구조적 유사도에 기반한 신뢰성 있는 루프 검출을 가능하게 하였고, 고해상도 센서, 키포인트 추출, 또는 학습 기반 전처리에 의존하지 않으면서 전역적으로 일관된 지도를 구축할 수 있도록 하였다.

    이러한 표현 중심의 기반 위에서, 위치 추정 단계에서는 자연어 설명과 공간 지도 데이터 간의 표현 불일치 문제가 다루어졌다. 텍스트 입력과 OpenStreetMap 데이터는 통합된 씬 그래프 표현으로 정규화되었으며, 이를 통해 공통 임베딩 공간에서 이종 모달 간 비교가 가능해졌다. 이러한 정규화 과정은 지도 정보가 불완전하거나 최신 상태가 아니거나 원시 센서 관측과의 상관성이 약한 환경에서도 신뢰성 있는 텍스트 기반 장소 인식을 가능하게 하였다.

    경로 계획 단계에서는 언어적 의도를 표현 학습의 대상으로 확장하였다. 자연어 명령은 암묵적인 선호와 제약 조건을 추출하도록 해석되었으며, 이는 그래프 기반 경로 계획기의 가중치 함수에 직접적으로 인코딩되었다. 이러한 의도 인지 표현을 OpenStreetMap 기반 경로 탐색에 통합함으로써, 생성된 경로는 최단 경로 최적화를 넘어 사용자 의도, 안전성, 그리고 상황적 맥락을 보다 잘 반영하도록 하였다.

    종합적으로, 본 학위논문은 환경적 제약 하에서 발생하는 자율주행 문제를 지도 작성, 위치 추정, 그리고 경로 계획 전반에 걸친 과제 지향적 표현 구성의 문제로 체계화하였다. 제안된 프레임워크는 기하적, 이종 모달, 그리고 언어적 추론을 통합적으로 연결함으로써 인간 중심적이며 강건한 자율주행의 가능성을 제시하였으며, 불확실성과 다양성, 그리고 대규모 특성을 갖는 실제 스마트시티 환경에서의 신뢰성 있는 자율 이동 시스템 구현 가능성을 입증하였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Autonomous navigation in smart-city environments demands robust operation under persistent environmental constraints, where sensing is imperfect, map information is inaccurate or incomplete, and user instructions are diverse, unstructured, and ambiguous. These conditions arise from extreme weather, limited visibility, signal attenuation, and sensor degradation, and they fundamentally challenge the reliability of autonomous mobility systems. A central difficulty lies in transforming heterogeneous and imperfect inputs, including spatial, cross-modal, and linguistic information, into task-relevant representations that support reliable navigation.

    This dissertation addresses this challenge from a unified representation-driven perspective, viewing autonomous navigation as a sequential pipeline consisting of mapping, localization, and path planning. Rather than treating failures at each stage independently, it identifies a common underlying problem: conventional representations are brittle under environmental constraints and fail to bridge geometric degradation, cross-modal inconsistency, and underutilized linguistic information. To overcome this limitation, the dissertation proposes an integrated framework that constructs task-driven representations tailored to each stage of the navigation pipeline while maintaining conceptual coherence across stages.

    At the mapping level, the framework focuses on environments where geometric observability is severely limited. Sparse and noisy point clouds are enriched through structural bootstrapping, which extracts complementary geometric, normal, and curvature statistics to form rotation-invariant structural representations. These representations enable reliable loop detection based on structural similarity, supporting globally consistent mapping without reliance on high-resolution sensors, keypoint extraction, or learning-based preprocessing.

    Building upon this representation-centric foundation, the localization stage addresses the misalignment between natural language descriptions and spatial map data. Textual inputs and OpenStreetMap data are normalized into unified scene graph representations, allowing cross-modal comparison in a shared embedding space. This canonicalization enables robust text-based place recognition even when map data are incomplete, outdated, or weakly correlated with raw sensor observations.

    At the path-planning stage, the framework extends representation learning to linguistic intent. Natural-language instructions are interpreted to extract implicit preferences and constraints, which are encoded directly into the weight function of a graph-based planner. By integrating these intent-aware representations into OpenStreetMap-based routing, the resulting paths move beyond shortest-distance optimization and better align with human intent, safety considerations, and situational context.

    Together, these contributions demonstrate that autonomous navigation under environmental constraints can be systematically addressed by constructing task-driven representations that unify geometric, cross-modal, and linguistic reasoning across mapping, localization, and path planning. The proposed framework advances human-centered and robust autonomous navigation, highlighting the potential for reliable deployment in real-world smart-city environments characterized by uncertainty, diversity, and scale.
    번역하기

    Autonomous navigation in smart-city environments demands robust operation under persistent environmental constraints, where sensing is imperfect, map information is inaccurate or incomplete, and user instructions are diverse, unstructured, and ambiguo...

    Autonomous navigation in smart-city environments demands robust operation under persistent environmental constraints, where sensing is imperfect, map information is inaccurate or incomplete, and user instructions are diverse, unstructured, and ambiguous. These conditions arise from extreme weather, limited visibility, signal attenuation, and sensor degradation, and they fundamentally challenge the reliability of autonomous mobility systems. A central difficulty lies in transforming heterogeneous and imperfect inputs, including spatial, cross-modal, and linguistic information, into task-relevant representations that support reliable navigation.

    This dissertation addresses this challenge from a unified representation-driven perspective, viewing autonomous navigation as a sequential pipeline consisting of mapping, localization, and path planning. Rather than treating failures at each stage independently, it identifies a common underlying problem: conventional representations are brittle under environmental constraints and fail to bridge geometric degradation, cross-modal inconsistency, and underutilized linguistic information. To overcome this limitation, the dissertation proposes an integrated framework that constructs task-driven representations tailored to each stage of the navigation pipeline while maintaining conceptual coherence across stages.

    At the mapping level, the framework focuses on environments where geometric observability is severely limited. Sparse and noisy point clouds are enriched through structural bootstrapping, which extracts complementary geometric, normal, and curvature statistics to form rotation-invariant structural representations. These representations enable reliable loop detection based on structural similarity, supporting globally consistent mapping without reliance on high-resolution sensors, keypoint extraction, or learning-based preprocessing.

    Building upon this representation-centric foundation, the localization stage addresses the misalignment between natural language descriptions and spatial map data. Textual inputs and OpenStreetMap data are normalized into unified scene graph representations, allowing cross-modal comparison in a shared embedding space. This canonicalization enables robust text-based place recognition even when map data are incomplete, outdated, or weakly correlated with raw sensor observations.

    At the path-planning stage, the framework extends representation learning to linguistic intent. Natural-language instructions are interpreted to extract implicit preferences and constraints, which are encoded directly into the weight function of a graph-based planner. By integrating these intent-aware representations into OpenStreetMap-based routing, the resulting paths move beyond shortest-distance optimization and better align with human intent, safety considerations, and situational context.

    Together, these contributions demonstrate that autonomous navigation under environmental constraints can be systematically addressed by constructing task-driven representations that unify geometric, cross-modal, and linguistic reasoning across mapping, localization, and path planning. The proposed framework advances human-centered and robust autonomous navigation, highlighting the potential for reliable deployment in real-world smart-city environments characterized by uncertainty, diversity, and scale.

    더보기

    목차 (Table of Contents)

    • 1 Introduction 1
    • 1 1 Autonomous Navigation in Smart Cities 1
    • 1 2 Challenges in Autonomous Navigation for Smart Cities 3
    • 1 3 Problem Definition 4
    • 1 4 Approach Overview 6
    • 1 Introduction 1
    • 1 1 Autonomous Navigation in Smart Cities 1
    • 1 2 Challenges in Autonomous Navigation for Smart Cities 3
    • 1 3 Problem Definition 4
    • 1 4 Approach Overview 6
    • 1 5 Contributions and Outline of the Dissertation 7
    • 2 Deficient Geometric Information: Structural Bootstrapping 10
    • 2 1 Introduction 10
    • 2 2 Related Works 14
    • 2 3 Problem Definition 18
    • 2 4 Proposed Methods 19
    • 2 4 1 Dead Reckoning 19
    • 2 4 2 Point Cloud Processing 20
    • 2 4 3 Feature Maps 21
    • 2 4 4 Loop Detection 25
    • 2 5 Experiments 26
    • 2 5 1 Research Questions 26
    • 2 5 2 Details 26
    • 2 5 3 Evaluation 27
    • 2 6 Conclusion 33
    • 3 Misaligned Cross-modal Information: Cross-Modal Canonicalization 35
    • 3 1 Introduction 35
    • 3 2 Related Works 37
    • 3 2 1 Point cloud map-based 37
    • 3 2 2 Scene graph-based 40
    • 3 2 3 OSM-based 41
    • 3 3 Problem Definition 42
    • 3 4 Methods 44
    • 3 4 1 Scene graph generation 44
    • 3 4 2 Similar scene graph candidates extraction 49
    • 3 4 3 Scene graph retrieval and scene selection 50
    • 3 5 Experiments 53
    • 3 5 1 Baselines and metrics 53
    • 3 5 2 Implementation 54
    • 3 5 3 Research questions 54
    • 3 5 4 Evaluation 55
    • 3 6 Conclusion 67
    • 4 Underutilized Linguistic Information: Foundation-driven Enrichment 73
    • 4 1 Introduction 73
    • 4 2 Related Works 76
    • 4 2 1 Pose-Goal Navigation 76
    • 4 2 2 Landmark-based Navigation 78
    • 4 2 3 Language-based Navigation 79
    • 4 3 System Overview 81
    • 4 4 Problem Definition 82
    • 4 5 Methods 83
    • 4 5 1 Graph Loader and Feature Extractor 83
    • 4 5 2 Intent Classifier 85
    • 4 5 3 Path Weighter 87
    • 4 5 4 Path Generator 87
    • 4 5 5 Path Evaluator 87
    • 4 5 6 Navigation 88
    • 4 6 Experiments 89
    • 4 6 1 Research Questions 89
    • 4 6 2 Experimental Setup 89
    • 4 6 3 Evaluation 91
    • 4 7 Conclusion 96
    • 5 Conclusion and Future Work 110
    • Abstract (In Korean) 130
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼