RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Learning Complete Scene Geometry from Partial Views Using Shape Primitives and Generative Models = 형상 원소와 생성 모델을 이용한 부분적 관측으로부터의 전체 환경 기하 복원 학습

    한글로보기

    https://www.riss.kr/link?id=T17452117

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    For robots to interact effectively with objects in their environment, it is essential to obtain three-dimensional information about those objects -- specifically, their shapes and positions. Such information is typically obtained from data provided by vision sensors, most commonly RGB images, sometimes accompanied by depth information.
    In recent years, fueled by rapid advances in machine learning, substantial progress has been made in processing visual data. Nevertheless, inferring accurate three-dimensional geometry and spatial configuration of objects solely from sparse and partially observed RGB images remains a highly challenging problem. Ensuring that machine learning–based perception methods function robustly and practically in real-world environments is an even greater challenge.

    This thesis presents learning-based methodologies that, given sparse and partial RGB observations, can simultaneously identify individual objects in a scene and accurately reconstruct their three-dimensional shapes and positions. The proposed methods are specifically designed for perceptual robustness in real-world conditions, spanning shape primitive based approaches to generative modeling. Through extensive real-world manipulation experiments, this work demonstrates that the proposed methods retain strong performance beyond simulation. The central contribution of this thesis lies in developing robust 3D object shape perception methods from incomplete RGB observations that generalize reliably to real-world scenarios.

    The first contribution is \textit{T$^2$SQNet} (Transparent Tableware SuperQuadric Network), a model that predicts low-dimensional geometric representations of transparent tableware objects using an extended deformable superquadrics. Leveraging the representational power of these primitives, the model captures the wide variety and complexity of real-world tableware geometries. Moreover, by relying solely on object mask images, the architecture minimizes the distribution gap between simulated and real-world data, thereby achieving strong real-world performance even when trained exclusively in simulation.
    As a byproduct and contribution of independent interest, we also present {\it TablewareNet}, a publicly available toolkit for generating diverse datasets of transparent tableware with varying shapes and sizes, constructed using the proposed extended deformable superquadric representation. Experiments show that T$^2$SQNet, when trained on data generated by TablewareNet, outperforms existing methods for transparent object perception and proves effective in robotic manipulation tasks such as grasping and target retrieval.

    Also, this thesis introduces DreamGrasp, a method that leverages pretrained generative image models trained on large-scale realistic datasets to infer occluded and unobserved regions of an environment. By combining coarse 3D reconstruction, contrastive-learning–based object separation, and text-guided per-object refinement, DreamGrasp overcomes the limitations of existing approaches and enables robust 3D reconstruction in cluttered, multi-object environments. Experimental results demonstrate that DreamGrasp not only reconstructs accurate object geometries but also achieves high success rates in sequential grasping and target retrieval tasks.
    번역하기

    For robots to interact effectively with objects in their environment, it is essential to obtain three-dimensional information about those objects -- specifically, their shapes and positions. Such information is typically obtained from data provided by...

    For robots to interact effectively with objects in their environment, it is essential to obtain three-dimensional information about those objects -- specifically, their shapes and positions. Such information is typically obtained from data provided by vision sensors, most commonly RGB images, sometimes accompanied by depth information.
    In recent years, fueled by rapid advances in machine learning, substantial progress has been made in processing visual data. Nevertheless, inferring accurate three-dimensional geometry and spatial configuration of objects solely from sparse and partially observed RGB images remains a highly challenging problem. Ensuring that machine learning–based perception methods function robustly and practically in real-world environments is an even greater challenge.

    This thesis presents learning-based methodologies that, given sparse and partial RGB observations, can simultaneously identify individual objects in a scene and accurately reconstruct their three-dimensional shapes and positions. The proposed methods are specifically designed for perceptual robustness in real-world conditions, spanning shape primitive based approaches to generative modeling. Through extensive real-world manipulation experiments, this work demonstrates that the proposed methods retain strong performance beyond simulation. The central contribution of this thesis lies in developing robust 3D object shape perception methods from incomplete RGB observations that generalize reliably to real-world scenarios.

    The first contribution is \textit{T$^2$SQNet} (Transparent Tableware SuperQuadric Network), a model that predicts low-dimensional geometric representations of transparent tableware objects using an extended deformable superquadrics. Leveraging the representational power of these primitives, the model captures the wide variety and complexity of real-world tableware geometries. Moreover, by relying solely on object mask images, the architecture minimizes the distribution gap between simulated and real-world data, thereby achieving strong real-world performance even when trained exclusively in simulation.
    As a byproduct and contribution of independent interest, we also present {\it TablewareNet}, a publicly available toolkit for generating diverse datasets of transparent tableware with varying shapes and sizes, constructed using the proposed extended deformable superquadric representation. Experiments show that T$^2$SQNet, when trained on data generated by TablewareNet, outperforms existing methods for transparent object perception and proves effective in robotic manipulation tasks such as grasping and target retrieval.

    Also, this thesis introduces DreamGrasp, a method that leverages pretrained generative image models trained on large-scale realistic datasets to infer occluded and unobserved regions of an environment. By combining coarse 3D reconstruction, contrastive-learning–based object separation, and text-guided per-object refinement, DreamGrasp overcomes the limitations of existing approaches and enables robust 3D reconstruction in cluttered, multi-object environments. Experimental results demonstrate that DreamGrasp not only reconstructs accurate object geometries but also achieves high success rates in sequential grasping and target retrieval tasks.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    로봇이 환경에 있는 물체들과 상호작용하기 위해서는, 우선 공간 내에 존재하는 물체들의 3차원 정보들, 즉 모양과 위치를 알아내야 한다.
    그리고 이러한 정보들은 때때로 깊이 정보와 함께 주로 RGB 이미지와 같은 비전 센서로부터 얻어지는 데이터의 형태로 주어진다.
    최근, 머신 러닝 기술의 눈부신 발전에 힘입어 이러한 비전 데이터를 처리하는 기술에 많은 발전이 있었지만, 부분적으로 관측된 RGB 이미지들만으로 물체들의 3차원 형상 및 위치를 파악하는 문제는 여전히 도전적인 문제이다.
    특히, 이러한 머신 러닝 기반의 인지 방법론을 실제 환경에서 실용적으로 강건하게 작동하도록 만드는 일은 더욱 어려운 문제이다.

    본 학위 논문에서는 희박하고 부분적으로 관측된 RGB 이미지들이 주어졌을 때, 환경에 존재하는 물체들을 개별로 식별함과 동시에 각 물체들의 위치 및 모양을 정확히 복구하는 러닝 기반 방법론들을 제시한다.
    특히, 본 논문에서 제안하는 방법론들은 형상 원소부터 생성 모델의 활용에 이르기까지 여러 방법들을 통해 현실 세계에서도 인지 능력을 보장할 수 있도록 설계되어있다.
    실제 세계에서의 물체 조작 실험들을 통해, 본 논문에서 제안하는 방법론들은 현실 세계에서도 그 성능을 보장할 수 있음을 입증한다. 본 학위 논문의 주요 기여는 이와 같이 현실 세계에서도 강건하게 작동하는 부분적 RGB 이미지 관측 기반의 3차원 물체 형상 인지 방법들을 제안하는데 있다.

    우선, 본 논문에서는 확장된 변형 가능한 슈퍼쿼드릭이라 불리우는 형상 원소를 이용하여 투명한 식기류들의 물체별 저차원 기하학적 표현를 예측하는 T$^2$SQNet (Transparent Tableware SuperQuadric Network)모델을 제안한다. 이 모델은 형상 원소의 높은 표현력을 이용하여 현실 세계의 다양하고 복잡한 식기류 형태를 표현할 수 있으며, 물체 마스크 이미지에만 기반한 네트워크 구조를 통해 시뮬레이션과 현실 세계에서의 입력 데이터 분포 차이를 최소화하여 시뮬레이션 데이터의 학습만으로 현실 세계에서의 성능을 보장할 수 있다.
    또한 이와 별개이면서 추가적인 기여로써, 본 논문에서는 제안한 확장된 변형 가능한 슈퍼쿼드릭을 이용하여 다양한 모양과 크기의 투명 식기류 데이터셋을 생성할 수 있는 공개 툴셋인 TablewareNet을 제안한다.
    실험을 통해 TablewareNet을 이용하여 학습한 T$^2$SQNet은 투명 물체를 인지함에 있어 기존의 방법론들보다 우수한 성능을 보였으며, 물체 파지 및 목표 물체 회수 작업을 통해 로봇 물체 조작에도 효과적으로 사용될 수 있음을 입증하였다.

    또한, 본 논문에서는 현실적인 대량의 데이터셋을 통해 미리 학습된 이미지 생성 모델의 능력을 이용하여, 주어진 환경의 관측되지 않은 가려진 부분을 추론하는 DreamGrasp를 제안한다.
    개략적 3차원 복원과 대조 학습을 이용한 물체 분리, 그리고 텍스트 기반의 각 물체별 정제 과정을 통해, DreamGrasp는 복잡하고 여러 물체가 있는 환경에서 기존 방법론들의 한계점을 우회하고 강건한 3차원 복원을 가능하게 한다.
    실험을 통해, DreamGrasp는 정확한 물체들의 기하학적 형태를 복원할 뿐만 아니라, 순차적인 물체 파지 및 목표 물체 회수에서도 높은 성공률을 보인다.
    번역하기

    로봇이 환경에 있는 물체들과 상호작용하기 위해서는, 우선 공간 내에 존재하는 물체들의 3차원 정보들, 즉 모양과 위치를 알아내야 한다. 그리고 이러한 정보들은 때때로 깊이 정보와 함께 ...

    로봇이 환경에 있는 물체들과 상호작용하기 위해서는, 우선 공간 내에 존재하는 물체들의 3차원 정보들, 즉 모양과 위치를 알아내야 한다.
    그리고 이러한 정보들은 때때로 깊이 정보와 함께 주로 RGB 이미지와 같은 비전 센서로부터 얻어지는 데이터의 형태로 주어진다.
    최근, 머신 러닝 기술의 눈부신 발전에 힘입어 이러한 비전 데이터를 처리하는 기술에 많은 발전이 있었지만, 부분적으로 관측된 RGB 이미지들만으로 물체들의 3차원 형상 및 위치를 파악하는 문제는 여전히 도전적인 문제이다.
    특히, 이러한 머신 러닝 기반의 인지 방법론을 실제 환경에서 실용적으로 강건하게 작동하도록 만드는 일은 더욱 어려운 문제이다.

    본 학위 논문에서는 희박하고 부분적으로 관측된 RGB 이미지들이 주어졌을 때, 환경에 존재하는 물체들을 개별로 식별함과 동시에 각 물체들의 위치 및 모양을 정확히 복구하는 러닝 기반 방법론들을 제시한다.
    특히, 본 논문에서 제안하는 방법론들은 형상 원소부터 생성 모델의 활용에 이르기까지 여러 방법들을 통해 현실 세계에서도 인지 능력을 보장할 수 있도록 설계되어있다.
    실제 세계에서의 물체 조작 실험들을 통해, 본 논문에서 제안하는 방법론들은 현실 세계에서도 그 성능을 보장할 수 있음을 입증한다. 본 학위 논문의 주요 기여는 이와 같이 현실 세계에서도 강건하게 작동하는 부분적 RGB 이미지 관측 기반의 3차원 물체 형상 인지 방법들을 제안하는데 있다.

    우선, 본 논문에서는 확장된 변형 가능한 슈퍼쿼드릭이라 불리우는 형상 원소를 이용하여 투명한 식기류들의 물체별 저차원 기하학적 표현를 예측하는 T$^2$SQNet (Transparent Tableware SuperQuadric Network)모델을 제안한다. 이 모델은 형상 원소의 높은 표현력을 이용하여 현실 세계의 다양하고 복잡한 식기류 형태를 표현할 수 있으며, 물체 마스크 이미지에만 기반한 네트워크 구조를 통해 시뮬레이션과 현실 세계에서의 입력 데이터 분포 차이를 최소화하여 시뮬레이션 데이터의 학습만으로 현실 세계에서의 성능을 보장할 수 있다.
    또한 이와 별개이면서 추가적인 기여로써, 본 논문에서는 제안한 확장된 변형 가능한 슈퍼쿼드릭을 이용하여 다양한 모양과 크기의 투명 식기류 데이터셋을 생성할 수 있는 공개 툴셋인 TablewareNet을 제안한다.
    실험을 통해 TablewareNet을 이용하여 학습한 T$^2$SQNet은 투명 물체를 인지함에 있어 기존의 방법론들보다 우수한 성능을 보였으며, 물체 파지 및 목표 물체 회수 작업을 통해 로봇 물체 조작에도 효과적으로 사용될 수 있음을 입증하였다.

    또한, 본 논문에서는 현실적인 대량의 데이터셋을 통해 미리 학습된 이미지 생성 모델의 능력을 이용하여, 주어진 환경의 관측되지 않은 가려진 부분을 추론하는 DreamGrasp를 제안한다.
    개략적 3차원 복원과 대조 학습을 이용한 물체 분리, 그리고 텍스트 기반의 각 물체별 정제 과정을 통해, DreamGrasp는 복잡하고 여러 물체가 있는 환경에서 기존 방법론들의 한계점을 우회하고 강건한 3차원 복원을 가능하게 한다.
    실험을 통해, DreamGrasp는 정확한 물체들의 기하학적 형태를 복원할 뿐만 아니라, 순차적인 물체 파지 및 목표 물체 회수에서도 높은 성공률을 보인다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • List of Tables ix
    • List of Figures xi
    • 1 Introduction 1
    • 1.1 Learning 3D Object Geometry for Real-World Object Manipulation 1
    • Abstract i
    • List of Tables ix
    • List of Figures xi
    • 1 Introduction 1
    • 1.1 Learning 3D Object Geometry for Real-World Object Manipulation 1
    • 1.2 Robust 3D Recognition from Partial-View RGB Images 3
    • 1.3 Contribution 4
    • 1.3.1 A Novel Transparent Object Recognition Model using Shape Primitives 4
    • 1.3.2 An Object Recognition Pipeline Leveraging Score Distillation Sampling 5
    • 1.4 Organization 6
    • 2 Preliminaries 9
    • 2.1 Introduction 9
    • 2.2 Deformable Superquadrics 10
    • 2.3 3D Object Generation with 2D Diffusion Models 14
    • 2.3.1 Diffusion Models 14
    • 2.3.2 Score Distillation Sampling 15
    • 2.3.3 Single-view 3D Generation 15
    • 3 T2SQNet: Transparent Tableware Superquadric Network 17
    • 3.1 Introduction 17
    • 3.2 Related Works 19
    • 3.2.1 NeRF-based Transparent Objects Recognition for Manipulation 19
    • 3.2.2 Learning-based Transparent Objects Recognition for Manipulation 20
    • 3.3 TablewareNet: Dataset for Cluttered Transparent Tableware 22
    • 3.3.1 Tableware Templates 22
    • 3.3.2 Synthetic Dataset Generation 24
    • 3.4 T2SQNet 27
    • 3.4.1 Model Architecture 30
    • 3.4.2 Training Process 33
    • 3.5 Geometry-Aware Object Manipulation with T2SQNet 37
    • 3.6 Experiments 39
    • 3.6.1 Transparent Object Recognition Performance 41
    • 3.6.2 Object Manipulation Performance 43
    • 3.6.3 Limitations and Future Directions 48
    • 3.7 Conclusion 49
    • 4 DreamGrasp: Multi-Object 3D Recognition Leveraging 2D Diffusion Models 51
    • 4.1 Introduction 51
    • 4.2 Related Works 53
    • 4.2.1 Robotic Object Manipulation with Differentiable 3D Representations 53
    • 4.2.2 3D Model Generation with Score Distillation Sampling 54
    • 4.3 Review: Distilling Image Generative Models for 3D Reconstruction 55
    • 4.4 DreamGrasp 58
    • 4.4.1 Coarse Scene Reconstruction Stage 58
    • 4.4.2 Instance-wise Refining Stage 63
    • 4.5 Object Manipulation with DreamGrasp 67
    • 4.6 Experiments 69
    • 4.6.1 Recognition Performance 70
    • 4.6.2 Object Manipulation Performance 71
    • 4.7 Limitations and Future Work 79
    • 4.8 Conclusion 80
    • 5 Conclusion 81
    • 5.1 Summary 81
    • 5.2 Future Work 84
    • 5.3 Concluding Remark 87
    • A Appendix: T2SQNet 89
    • A.1 Implementation Details for Our Methods 89
    • A.1.1 Details for TablewareNet Objects 89
    • A.1.2 Examples of Class Supplementation Using Deformable Superquadrics 96
    • A.1.3 Details for Model Architecture and Training Process 98
    • A.1.4 Details for Geometry-Aware Object Manipulation 98
    • A.2 Experimental Details 102
    • A.2.1 Additional Details for Recognition Experiments 102
    • A.2.2 Additional Details for Sequential Declutter Experiment 105
    • A.2.3 Additional Details for Target Retrieval Experiment 106
    • A.3 Additional Experimental Results 107
    • A.3.1 Additional Results for Recognition Experiments 107
    • A.3.2 Additional Results for Sequential Declutter Experiments 118
    • A.3.3 Additional Results for Target Retrieval Experiments 120
    • A.3.4 Comparison of T2SQNet with an End-to-End Method 123
    • B Appendix: DreamGrasp 125
    • B.1 Implementation Details of DreamGrasp 125
    • B.1.1 Coarse Stage 125
    • B.1.2 Refinement Stage 126
    • B.2 Details on Object Manipulation with DreamGrasp 127
    • B.2.1 Details on Grasp Pose Generation 127
    • B.2.2 Details on Collision Detection 127
    • B.2.3 Details on Sequential Declutter 128
    • B.2.4 Details on Target Retrieval 129
    • B.3 Details on Experiments and Additional Results 130
    • B.3.1 Details on Experimental Settings 130
    • B.3.2 Exteneded and Additional Recognition Experiments Results 130
    • Bibliography 137
    • Abstract 152
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼