RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Vision-Based Hierarchical Policy Learning for Robotic Object Rearrangement = 물체 재배열을 위한 시각 기반 계층 정책 학습

    한글로보기

    https://www.riss.kr/link?id=T17451168

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Robotic object rearrangement is a core capability for manipulation in human-centric environments, where robots must interact with diverse objects under clutter, occlusions, and dynamic scene changes.
    These settings demand decision making under partial observability, combinatorial multi-object interactions, and long-horizon objectives, while respecting geometric, kinematic, and safety constraints.
    This dissertation develops a unified framework for vision-based hierarchical policy learning with learning-based action primitives, spanning (i) goal-conditioned tabletop rearrangement with explicit target configurations, (ii) goal-free tabletop tidying under implicit organizational preferences, and (iii) indoor humanoid loco-manipulation that requires coordinating navigation and interaction in large 3D environments.

    A central thesis of this work is that directly learning end-to-end whole-body control is often inefficient and brittle for rearrangement, particularly for high-DoF robots and contact-rich tasks.
    Instead, we construct learning-based action primitives that encapsulate meaningful interaction skills and learn high-level decision-making policies that select and parameterize these primitives from visual observations.
    Building on this perspective, the dissertation presents four instantiations of hierarchical action primitives and the corresponding learning algorithms.

    First, we study goal-conditioned rearrangement using non-prehensile pushing primitives.
    We introduce a vision-based decomposed Q-learning framework that parameterizes pushes in image space and learns value functions for selecting effective push actions under partial observability.
    This formulation improves data efficiency by structuring the action space around a manipulation primitive rather than raw joint control.

    Second, we improve robustness and generalization by incorporating geometric scene structure.
    We propose signed distance field (SDF)-based scene representations and object-centric scene graphs that encode object geometry and spatial relations.
    Based on this representation, we develop an SDF-based push primitive and a Q-learning method that generalizes across diverse object configurations and varying numbers of objects in cluttered tabletop environments.

    Third, we address goal-free tabletop tidying, where specifying a single target arrangement is impractical and success is defined by an implicit notion of organization.
    We construct the Tabletop Tidying Up (TTU) dataset and learn the Tidiness Score, a visual metric that quantifies the degree of tidiness across diverse clutter patterns and environments.
    We then define grasp-parameterized pick-and-place primitives using predicted 6-DoF grasping points, train a tidying policy with offline reinforcement learning, and integrate the learned policy and tidiness score into a Monte Carlo Tree Search (MCTS) planner to enable long-horizon decision making for tidying.

    Finally, we extend the action-primitive hierarchy to indoor humanoid loco-manipulation.
    We present a hierarchical whole-body control framework that couples a vision-language-action (VLA) policy for upper-body manipulation with a separate locomotion policy for stable lower-body control.
    The high-level policy outputs high-level locomotion commands (e.g., linear velocity and target yaw) and manipulation actions, while a mode gate mediates between locomotion and manipulation-dominant behaviors across different phases of a task, enabling scalable coordination for large-space rearrangement.

    Overall, this dissertation demonstrates that vision-based hierarchical policy learning with learning-based action primitives-combining structured representations, value-based learning, offline policy learning, and planning-provides a principled and scalable approach to robust object rearrangement across tasks, embodiments, and unstructured environments.
    번역하기

    Robotic object rearrangement is a core capability for manipulation in human-centric environments, where robots must interact with diverse objects under clutter, occlusions, and dynamic scene changes. These settings demand decision making under partial...

    Robotic object rearrangement is a core capability for manipulation in human-centric environments, where robots must interact with diverse objects under clutter, occlusions, and dynamic scene changes.
    These settings demand decision making under partial observability, combinatorial multi-object interactions, and long-horizon objectives, while respecting geometric, kinematic, and safety constraints.
    This dissertation develops a unified framework for vision-based hierarchical policy learning with learning-based action primitives, spanning (i) goal-conditioned tabletop rearrangement with explicit target configurations, (ii) goal-free tabletop tidying under implicit organizational preferences, and (iii) indoor humanoid loco-manipulation that requires coordinating navigation and interaction in large 3D environments.

    A central thesis of this work is that directly learning end-to-end whole-body control is often inefficient and brittle for rearrangement, particularly for high-DoF robots and contact-rich tasks.
    Instead, we construct learning-based action primitives that encapsulate meaningful interaction skills and learn high-level decision-making policies that select and parameterize these primitives from visual observations.
    Building on this perspective, the dissertation presents four instantiations of hierarchical action primitives and the corresponding learning algorithms.

    First, we study goal-conditioned rearrangement using non-prehensile pushing primitives.
    We introduce a vision-based decomposed Q-learning framework that parameterizes pushes in image space and learns value functions for selecting effective push actions under partial observability.
    This formulation improves data efficiency by structuring the action space around a manipulation primitive rather than raw joint control.

    Second, we improve robustness and generalization by incorporating geometric scene structure.
    We propose signed distance field (SDF)-based scene representations and object-centric scene graphs that encode object geometry and spatial relations.
    Based on this representation, we develop an SDF-based push primitive and a Q-learning method that generalizes across diverse object configurations and varying numbers of objects in cluttered tabletop environments.

    Third, we address goal-free tabletop tidying, where specifying a single target arrangement is impractical and success is defined by an implicit notion of organization.
    We construct the Tabletop Tidying Up (TTU) dataset and learn the Tidiness Score, a visual metric that quantifies the degree of tidiness across diverse clutter patterns and environments.
    We then define grasp-parameterized pick-and-place primitives using predicted 6-DoF grasping points, train a tidying policy with offline reinforcement learning, and integrate the learned policy and tidiness score into a Monte Carlo Tree Search (MCTS) planner to enable long-horizon decision making for tidying.

    Finally, we extend the action-primitive hierarchy to indoor humanoid loco-manipulation.
    We present a hierarchical whole-body control framework that couples a vision-language-action (VLA) policy for upper-body manipulation with a separate locomotion policy for stable lower-body control.
    The high-level policy outputs high-level locomotion commands (e.g., linear velocity and target yaw) and manipulation actions, while a mode gate mediates between locomotion and manipulation-dominant behaviors across different phases of a task, enabling scalable coordination for large-space rearrangement.

    Overall, this dissertation demonstrates that vision-based hierarchical policy learning with learning-based action primitives-combining structured representations, value-based learning, offline policy learning, and planning-provides a principled and scalable approach to robust object rearrangement across tasks, embodiments, and unstructured environments.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    로봇 물체 재배열은 인간 중심 환경에서의 매니퓰레이션에 핵심적인 능력으로 로봇은 혼잡하고 가려짐이 많은 상황과 동적인 장면 변화 속에서 다양한 물체와 상호작용해야 한다.
    이러한 상황에서는 부분 관측 하에서의 의사결정과 조합적으로 증가하는 다중 물체 상호작용, 그리고 장시간 목표를 요구하며 동시에 기하학적·운동학적·안전 제약을 만족해야 한다.
    본 논문에서는 학습 기반 Action Primitive를 활용하는 비전 기반 계층 정책 학습 프레임워크를 제안하며, 이를 (i) 명시적 목표 구성이 주어진 목표 기반 테이블 재배열, (ii) 명시적 목표 없이 정리 정도에 의해 정의되는 테이블 정돈, (iii) 3차원 실내 공간에서의 이동과 상호작용을 요구하는 휴머노이드 로코-매니퓰레이션으로 확장한다.

    본 논문의 핵심 논지는, 특히 고자유도 로봇과 접촉 상호작용이 많은 과제에서 End-To-End 방식의 전신 제어 학습이 비효율적이며 안정성이 취약하다는 점이다.
    대신, 의미 있는 상호작용 기술을 내포하는 학습 기반 행동 프리미티브를 구성하고 시각 관측으로부터 이러한 프리미티브를 선택·매개변수화하는 고수준 정책을 학습하는 계층 구조를 채택한다.
    이러한 관점을 기반으로 본 논문은 네 가지 형태의 계층적 행동 프리미티브와 그에 상응하는 학습 알고리즘을 제시한다.

    첫째, 비파지 밀기 동작을 통한 목표 조건형 재배열을 다룬다.
    이미지 공간에서 파라미터화된 Push Primitive를 기반으로 부분 관측 환경에서 효과적인 로봇 동작을 선택하기 위한 가치 함수를 학습하는 시각 기반 분해형 Q-함수 학습 프레임워크를 제안한다.
    이러한 구조화된 동작 표현은 직접적인 관절 제어 대신 조작 프리미티브를 중심으로 행동 공간을 구조화함으로써 데이터 효율성을 향상시킨다.

    둘째, 기하학적 장면 구조를 통합하여 강건성과 일반화를 개선한다.
    이를 위해 물체의 기하학적인 정보와 공간적 관계를 인코딩하는 거리부호함수 기반의 장면 표현과 객체 중심의 장면 그래프를 제안한다.
    이 표현에 기반하여 거리부호 함수 기반 Push Primitive에 따른 Q-함수 학습 방법을 제안하며, 이는 복잡한 테이블탑 환경에서 다양한 물체 배치와 서로 다른 물체 개수에 대해 일반화되는 성질을 갖는다.

    셋째, 명시적 목표 구성이 어려우며 정돈 여부가 암묵적 기준에 의해 정의되는 목표 비지정 테이블 정돈 문제를 다룬다.
    이를 위해 테이블탑 정리정돈 데이터셋을 구축하고 다양한 혼잡 패턴과 환경 전반에서 정리 정도를 정량화하는 시각적 지표인 정리 점수를 학습한다.
    이후 예측된 6-자유도 파지 지점을 기반으로 Pick-And-Place Primitive를 정의하고, 오프라인 강화학습을 통해 정리 정책을 학습한다.
    또한 학습된 정리 정책과 정리 점수를 몬테카를로 트리 탐색과 통합하여 정리 과제에서 장기 의사결정을 가능하게 한다.

    마지막으로 Action Primitive 기반 계층 구조를 실내 휴머노이드 로코-매니퓰레이션으로 확장한다.
    우리는 상체 조작을 위한 시각-언어-행동 정책과 안정적인 하체 제어를 위한 별도의 보행 정책을 결합한 계층적 전신 제어 프레임워크를 제시한다.
    고수준 정책은 고수준 보행 명령과 조작 행동을 출력하며, 모드 게이트는 과제의 단계에 따라 보행 중심 행동과 조작 중심 행동 간 전환을 조절하여 대규모 공간에서의 정리·재배열 과제를 수행할 수 있도록 한다.

    종합하면 본 논문은 구조화된 표현, 가치 기반 학습, 오프라인 정책 학습, 그리고 계획 기반 탐색을 결합한 시각 기반 계층적 정책 학습 및 학습 기반 Action Primitive가 다양한 과제, 다양한 로봇 형태, 비정형 환경에서 강인하고 확장성 있는 물체 재배열을 가능하게 하는 원리적 접근임을 보여준다.
    번역하기

    로봇 물체 재배열은 인간 중심 환경에서의 매니퓰레이션에 핵심적인 능력으로 로봇은 혼잡하고 가려짐이 많은 상황과 동적인 장면 변화 속에서 다양한 물체와 상호작용해야 한다. 이러한 ...

    로봇 물체 재배열은 인간 중심 환경에서의 매니퓰레이션에 핵심적인 능력으로 로봇은 혼잡하고 가려짐이 많은 상황과 동적인 장면 변화 속에서 다양한 물체와 상호작용해야 한다.
    이러한 상황에서는 부분 관측 하에서의 의사결정과 조합적으로 증가하는 다중 물체 상호작용, 그리고 장시간 목표를 요구하며 동시에 기하학적·운동학적·안전 제약을 만족해야 한다.
    본 논문에서는 학습 기반 Action Primitive를 활용하는 비전 기반 계층 정책 학습 프레임워크를 제안하며, 이를 (i) 명시적 목표 구성이 주어진 목표 기반 테이블 재배열, (ii) 명시적 목표 없이 정리 정도에 의해 정의되는 테이블 정돈, (iii) 3차원 실내 공간에서의 이동과 상호작용을 요구하는 휴머노이드 로코-매니퓰레이션으로 확장한다.

    본 논문의 핵심 논지는, 특히 고자유도 로봇과 접촉 상호작용이 많은 과제에서 End-To-End 방식의 전신 제어 학습이 비효율적이며 안정성이 취약하다는 점이다.
    대신, 의미 있는 상호작용 기술을 내포하는 학습 기반 행동 프리미티브를 구성하고 시각 관측으로부터 이러한 프리미티브를 선택·매개변수화하는 고수준 정책을 학습하는 계층 구조를 채택한다.
    이러한 관점을 기반으로 본 논문은 네 가지 형태의 계층적 행동 프리미티브와 그에 상응하는 학습 알고리즘을 제시한다.

    첫째, 비파지 밀기 동작을 통한 목표 조건형 재배열을 다룬다.
    이미지 공간에서 파라미터화된 Push Primitive를 기반으로 부분 관측 환경에서 효과적인 로봇 동작을 선택하기 위한 가치 함수를 학습하는 시각 기반 분해형 Q-함수 학습 프레임워크를 제안한다.
    이러한 구조화된 동작 표현은 직접적인 관절 제어 대신 조작 프리미티브를 중심으로 행동 공간을 구조화함으로써 데이터 효율성을 향상시킨다.

    둘째, 기하학적 장면 구조를 통합하여 강건성과 일반화를 개선한다.
    이를 위해 물체의 기하학적인 정보와 공간적 관계를 인코딩하는 거리부호함수 기반의 장면 표현과 객체 중심의 장면 그래프를 제안한다.
    이 표현에 기반하여 거리부호 함수 기반 Push Primitive에 따른 Q-함수 학습 방법을 제안하며, 이는 복잡한 테이블탑 환경에서 다양한 물체 배치와 서로 다른 물체 개수에 대해 일반화되는 성질을 갖는다.

    셋째, 명시적 목표 구성이 어려우며 정돈 여부가 암묵적 기준에 의해 정의되는 목표 비지정 테이블 정돈 문제를 다룬다.
    이를 위해 테이블탑 정리정돈 데이터셋을 구축하고 다양한 혼잡 패턴과 환경 전반에서 정리 정도를 정량화하는 시각적 지표인 정리 점수를 학습한다.
    이후 예측된 6-자유도 파지 지점을 기반으로 Pick-And-Place Primitive를 정의하고, 오프라인 강화학습을 통해 정리 정책을 학습한다.
    또한 학습된 정리 정책과 정리 점수를 몬테카를로 트리 탐색과 통합하여 정리 과제에서 장기 의사결정을 가능하게 한다.

    마지막으로 Action Primitive 기반 계층 구조를 실내 휴머노이드 로코-매니퓰레이션으로 확장한다.
    우리는 상체 조작을 위한 시각-언어-행동 정책과 안정적인 하체 제어를 위한 별도의 보행 정책을 결합한 계층적 전신 제어 프레임워크를 제시한다.
    고수준 정책은 고수준 보행 명령과 조작 행동을 출력하며, 모드 게이트는 과제의 단계에 따라 보행 중심 행동과 조작 중심 행동 간 전환을 조절하여 대규모 공간에서의 정리·재배열 과제를 수행할 수 있도록 한다.

    종합하면 본 논문은 구조화된 표현, 가치 기반 학습, 오프라인 정책 학습, 그리고 계획 기반 탐색을 결합한 시각 기반 계층적 정책 학습 및 학습 기반 Action Primitive가 다양한 과제, 다양한 로봇 형태, 비정형 환경에서 강인하고 확장성 있는 물체 재배열을 가능하게 하는 원리적 접근임을 보여준다.

    더보기

    목차 (Table of Contents)

    • 1 Introduction 1
    • 1.1 Motivation 1
    • 1.2 Organization of the Dissertation 6
    • 2 Background 11
    • 2.1 Sequential Decision-Making Formulations 12
    • 1 Introduction 1
    • 1.1 Motivation 1
    • 1.2 Organization of the Dissertation 6
    • 2 Background 11
    • 2.1 Sequential Decision-Making Formulations 12
    • 2.1.1 Markov Decision Process 12
    • 2.1.2 Partially Observable Markov Decision Process 13
    • 2.1.3 Multi-Objective Markov Decision Processes 14
    • 2.2 Reinforcement Learning and Data-Driven Control Methods 15
    • 2.2.1 Reinforcement Learning 15
    • 2.2.2 Deep Q-Learning 16
    • 2.2.3 Goal-Conditioned Deep Q-Learning 17
    • 2.2.4 Hindsight Experience Replay 18
    • 2.2.5 Imitation Learning 18
    • 2.2.6 Offline Reinforcement Learning 19
    • 2.2.7 Implicit Q-Learning 20
    • 2.3 Geometric and Graph-Based Scene Representations and Planning 21
    • 2.3.1 Signed Distance Field 21
    • 2.3.2 Fast Marching Method 22
    • 2.3.3 Graph Convolutional Networks 23
    • 2.3.4 Monte Carlo Tree Search 23
    • 2.4 Vision-Language and Vision-Language-Action Models 25
    • 2.4.1 Vision-Language Models 25
    • 2.4.2 Vision-Language-Action Models 25
    • 3 Vision-Based Decomposed Q-Learning for Non-Prehensile Object Rearrangement 27
    • 3.1 Decomposed Q-Learning for Non-Prehensile Rearrangement Problem 27
    • 3.1.1 Related Work 28
    • 3.1.2 Methods 30
    • 3.1.3 Deep Q-Networks 31
    • 3.1.4 Decomposed Q-Networks 31
    • 3.1.5 Experiments 33
    • 3.2 Scene Segmentation Reasoning for Multi-object Manipulation 34
    • 3.2.1 Problem Definition 37
    • 3.2.2 Methods 38
    • 3.2.3 Scene Segmentation Reasoning 38
    • 3.2.4 Q-learning 39
    • 3.2.5 Hindsight Experience Replay 41
    • 3.2.6 Experiments 41
    • 3.3 Summary 42
    • 4 Vision-Based Q-Learning with SDF Scene Graphs for Tabletop Rearrangement 45
    • 4.1 Motivation 46
    • 4.2 Related Work 49
    • 4.3 Problem Description 51
    • 4.4 Proposed Method 52
    • 4.4.1 SDF-based Scene Graph 54
    • 4.4.2 Graph Convolutional Network on SDF Nodes 56
    • 4.4.3 Training SDFGCN 57
    • 4.5 Experiments 59
    • 4.5.1 Experimental Setup 59
    • 4.5.2 Comparison to Baseline Methods 60
    • 4.5.3 Generalization to Different Number of Objects 62
    • 4.5.4 Real Robot Experiments 63
    • 4.6 Summary 67
    • 5 Vision-Based Hierarchical Planning for Tabletop Tidying Up 69
    • 5.1 Motivation 70
    • 5.2 Related Work 73
    • 5.3 Problem Formulation 74
    • 5.4 Tabletop Tidying Up Dataset 75
    • 5.5 Proposed Method 77
    • 5.5.1 Tidiness Discriminator and Tidying Policy 79
    • 5.5.2 Low-Level Planner 82
    • 5.5.3 Tidiness Score-Guided High-Level Planner 85
    • 5.6 Experiments 89
    • 5.6.1 Evaluation of Tidiness Discriminator 89
    • 5.6.2 Simulation Experiments 92
    • 5.6.3 Real Robot Experiments 96
    • 5.7 Summary 100
    • 6 Vision-Based Hierarchical VLA Policy Learning for Humanoid Loco-Manipulation 103
    • 6.1 Motivation 104
    • 6.2 Related Work 108
    • 6.2.1 Humanoid Whole-Body Control 108
    • 6.2.2 Vision-Language-Action Models 109
    • 6.3 Problem Formulation 111
    • 6.3.1 Proprioceptive State 111
    • 6.3.2 Hierarchical Policy for Whole-body Control 112
    • 6.4 Methods 113
    • 6.4.1 Hierarchical Whole-body Policy: VLA + AMO 114
    • 6.4.2 Two-Expert Gated Diffusion Policy in the VLA 116
    • 6.4.3 Training Objective and Responsibility-Based Routing 122
    • 6.5 Experiments 124
    • 6.5.1 Robot Platform and Simulation Environment 125
    • 6.5.2 Loco-Manipulation Tasks and Evaluation Protocol 126
    • 6.5.3 Data Collection and Dataset Construction 128
    • 6.5.4 Baselines and Ablations 130
    • 6.5.5 Simulation Results 132
    • 6.6 Summary 138
    • 6.7 Future Work 139
    • 7 Conclusion 143
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼