RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Toward Reliable Vision-Based Robotic Grasping in Transparency and Clutter = 투명 물체 및 복잡한 환경에서의 신뢰성 있는 비전 기반 로봇 파지 연구

    한글로보기

    https://www.riss.kr/link?id=T17313243

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Vision-based grasping focuses on determining successful grasp configurations using input from vision sensors.
    Typically, algorithms take RGB and depth images from modern RGB-D cameras and output where and how to grasp a target object in 3D space.
    With advances in grasping algorithms and vision sensors, vision-based grasping has become more reliable across a wide range of scenarios.
    Grasping models, enhanced by sophisticated engineering, can incorporate physically grounded biases—such as smoothness and center of gravity—to generate high-quality grasps based on object geometry.
    At the same time, vision sensors leveraging technologies like stereo vision, LiDAR, and infrared have become more capable and affordable, allowing accurate geometry capture for most objects.
    Moreover, the increasing availability of real-world datasets has significantly boosted performance in practical robotic applications.

    However, scenes containing transparency and clutter often fail to be correctly recognized by vision sensors, greatly reducing the reliability of vision-based grasping algorithms.
    First, transparent objects cause complex sensing failures in depth cameras due to its physical properties.
    Second, clutter, where multiple objects are in contact, results in occlusions and unobservable surfaces.
    These elements obstruct the acquisition of accurate geometry in a vision-based manner, leading to inaccurate grasps.

    In this thesis, I propose a reliable approach for acquiring scene geometry in environments with transparency and clutter.
    By leveraging information available from general pretrained vision modules, the method focuses on the aspect of generalization to various scenes.
    It extracts and utilizes mid-level representations such as masks and surface normals in spatially structured ways to achieve stable geometric reconstruction.
    First, I introduce a data-driven method for reliably obtaining instance masks in cluttered scenes containing transparent objects, along with a corresponding grasping algorithm.
    Second, I propose a novel 3D representation based on surface normals to effectively capture the geometry of transparent objects.
    Finally, I present an interactive perception pipeline that actively acquires instance-level 3D geometry in cluttered scenes with occlusions.
    Furthermore, experiments conducted on a real-world robotic platform demonstrate the potential for practical deployment in real-world scenarios.

    To enable robots to effectively replace human labor across diverse scenarios, consistent performance in variable conditions must be ensured.
    By addressing transparency and clutter—two major challenges for vision-based grasping—and proposing solutions built on general-purpose vision modules, this thesis aims to improve the stability and robustness of robotic manipulation.
    번역하기

    Vision-based grasping focuses on determining successful grasp configurations using input from vision sensors. Typically, algorithms take RGB and depth images from modern RGB-D cameras and output where and how to grasp a target object in 3D space. With...

    Vision-based grasping focuses on determining successful grasp configurations using input from vision sensors.
    Typically, algorithms take RGB and depth images from modern RGB-D cameras and output where and how to grasp a target object in 3D space.
    With advances in grasping algorithms and vision sensors, vision-based grasping has become more reliable across a wide range of scenarios.
    Grasping models, enhanced by sophisticated engineering, can incorporate physically grounded biases—such as smoothness and center of gravity—to generate high-quality grasps based on object geometry.
    At the same time, vision sensors leveraging technologies like stereo vision, LiDAR, and infrared have become more capable and affordable, allowing accurate geometry capture for most objects.
    Moreover, the increasing availability of real-world datasets has significantly boosted performance in practical robotic applications.

    However, scenes containing transparency and clutter often fail to be correctly recognized by vision sensors, greatly reducing the reliability of vision-based grasping algorithms.
    First, transparent objects cause complex sensing failures in depth cameras due to its physical properties.
    Second, clutter, where multiple objects are in contact, results in occlusions and unobservable surfaces.
    These elements obstruct the acquisition of accurate geometry in a vision-based manner, leading to inaccurate grasps.

    In this thesis, I propose a reliable approach for acquiring scene geometry in environments with transparency and clutter.
    By leveraging information available from general pretrained vision modules, the method focuses on the aspect of generalization to various scenes.
    It extracts and utilizes mid-level representations such as masks and surface normals in spatially structured ways to achieve stable geometric reconstruction.
    First, I introduce a data-driven method for reliably obtaining instance masks in cluttered scenes containing transparent objects, along with a corresponding grasping algorithm.
    Second, I propose a novel 3D representation based on surface normals to effectively capture the geometry of transparent objects.
    Finally, I present an interactive perception pipeline that actively acquires instance-level 3D geometry in cluttered scenes with occlusions.
    Furthermore, experiments conducted on a real-world robotic platform demonstrate the potential for practical deployment in real-world scenarios.

    To enable robots to effectively replace human labor across diverse scenarios, consistent performance in variable conditions must be ensured.
    By addressing transparency and clutter—two major challenges for vision-based grasping—and proposing solutions built on general-purpose vision modules, this thesis aims to improve the stability and robustness of robotic manipulation.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    비전 기반 파지는 로봇이 시각 센서로부터 얻은 정보를 바탕으로 성공적인 파지 구성을 결정하는 데 초점을 둔다. 일반적으로 이러한 알고리즘은 최신 RGB-D 카메라로부터 획득한 RGB 및 깊이 영상을 입력으로 받아, 3차원 공간에서 물체를 어떻게, 어디서 파지할지를 출력한다. 이 방식은 비정형 환경에서의 로봇 조작을 가능하게 하는 잠재력 덕분에 활발히 연구되어 왔다. 특히, 임의의 장면에서 처음 마주하는 물체를 다뤄야 할 때, 비전 기반 파지는 물체 정렬, 빈 정리 등과 같은 유의미한 작업 수행을 위한 핵심 구성 요소로 작용한다.

    최근 파지 알고리즘과 시각 센서의 발전으로, 비전 기반 파지는 다양한 상황에서도 점점 더 신뢰성 있는 성능을 보이고 있다. 최신 파지 모델은 정교한 공학적 설계를 바탕으로, 표면의 매끄러움이나 무게중심과 같은 물리 기반의 편향을 반영해 물체의 기하 정보를 활용한 고품질 파지를 생성할 수 있다. 동시에 스테레오 비전, LiDAR, 적외선 기반 기술을 활용한 시각 센서 역시 정밀도는 물론 가격 면에서도 개선되어 대부분의 물체에 대해 정확한 기하 추정을 가능하게 하고 있다. 또한, 실세계 기반 데이터셋의 확산은 실제 로봇 응용 분야에서의 성능 향상에 크게 기여하고 있다.

    그러나 투명 물체나 복잡하게 얽힌 장면(clutter)에서는 여전히 시각 센서가 올바른 인식을 하지 못해, 비전 기반 파지 알고리즘의 신뢰도를 크게 떨어뜨리는 문제가 발생한다. 첫째, 투명 물체는 물리적 특성상 깊이 카메라에서 복잡한 센싱 오류를 유발한다. 둘째, 여러 물체가 접촉해 있는 클러터 환경에서는 가려진 표면이나 관찰이 불가능한 영역이 많아져 정확한 기하 정보 획득이 어렵다. 이러한 요소들은 시각 기반의 정확한 기하 추정을 방해하며, 결과적으로 파지 실패로 이어질 수 있다.

    본 논문에서는 투명 물체 및 클러터 환경에서도 신뢰성 있는 기하 정보를 획득할 수 있는 방법을 제안한다. 일반 데이터로 학습된 사전 학습 비전 모듈의 출력을 활용함으로써, 특정 장면에 의존한 튜닝 없이도 안정적인 기하 재구성이 가능하도록 한다. 구체적으로, 마스크와 표면 법선과 같은 중간 수준의 표현을 공간적으로 구조화하여 활용함으로써 견고한 기하 정보를 추출한다. 또한, 실제 로봇 시스템을 활용한 실험을 통해 제안된 방법의 실용성과 현장 적용 가능성을 입증하였다.

    다양한 환경에서 로봇이 인간의 노동을 효과적으로 대체하기 위해서는, 환경 변화에 따른 성능 안정성이 반드시 보장되어야 한다. 본 논문은 비전 기반 파지의 대표적 난제인 투명성과 클러터 문제를 해결하고, 일반 비전 모듈을 기반으로 한 해법을 제안함으로써 로봇 조작의 안정성과 견고함 향상에 기여하고자 한다.
    번역하기

    비전 기반 파지는 로봇이 시각 센서로부터 얻은 정보를 바탕으로 성공적인 파지 구성을 결정하는 데 초점을 둔다. 일반적으로 이러한 알고리즘은 최신 RGB-D 카메라로부터 획득한 RGB 및 깊이 ...

    비전 기반 파지는 로봇이 시각 센서로부터 얻은 정보를 바탕으로 성공적인 파지 구성을 결정하는 데 초점을 둔다. 일반적으로 이러한 알고리즘은 최신 RGB-D 카메라로부터 획득한 RGB 및 깊이 영상을 입력으로 받아, 3차원 공간에서 물체를 어떻게, 어디서 파지할지를 출력한다. 이 방식은 비정형 환경에서의 로봇 조작을 가능하게 하는 잠재력 덕분에 활발히 연구되어 왔다. 특히, 임의의 장면에서 처음 마주하는 물체를 다뤄야 할 때, 비전 기반 파지는 물체 정렬, 빈 정리 등과 같은 유의미한 작업 수행을 위한 핵심 구성 요소로 작용한다.

    최근 파지 알고리즘과 시각 센서의 발전으로, 비전 기반 파지는 다양한 상황에서도 점점 더 신뢰성 있는 성능을 보이고 있다. 최신 파지 모델은 정교한 공학적 설계를 바탕으로, 표면의 매끄러움이나 무게중심과 같은 물리 기반의 편향을 반영해 물체의 기하 정보를 활용한 고품질 파지를 생성할 수 있다. 동시에 스테레오 비전, LiDAR, 적외선 기반 기술을 활용한 시각 센서 역시 정밀도는 물론 가격 면에서도 개선되어 대부분의 물체에 대해 정확한 기하 추정을 가능하게 하고 있다. 또한, 실세계 기반 데이터셋의 확산은 실제 로봇 응용 분야에서의 성능 향상에 크게 기여하고 있다.

    그러나 투명 물체나 복잡하게 얽힌 장면(clutter)에서는 여전히 시각 센서가 올바른 인식을 하지 못해, 비전 기반 파지 알고리즘의 신뢰도를 크게 떨어뜨리는 문제가 발생한다. 첫째, 투명 물체는 물리적 특성상 깊이 카메라에서 복잡한 센싱 오류를 유발한다. 둘째, 여러 물체가 접촉해 있는 클러터 환경에서는 가려진 표면이나 관찰이 불가능한 영역이 많아져 정확한 기하 정보 획득이 어렵다. 이러한 요소들은 시각 기반의 정확한 기하 추정을 방해하며, 결과적으로 파지 실패로 이어질 수 있다.

    본 논문에서는 투명 물체 및 클러터 환경에서도 신뢰성 있는 기하 정보를 획득할 수 있는 방법을 제안한다. 일반 데이터로 학습된 사전 학습 비전 모듈의 출력을 활용함으로써, 특정 장면에 의존한 튜닝 없이도 안정적인 기하 재구성이 가능하도록 한다. 구체적으로, 마스크와 표면 법선과 같은 중간 수준의 표현을 공간적으로 구조화하여 활용함으로써 견고한 기하 정보를 추출한다. 또한, 실제 로봇 시스템을 활용한 실험을 통해 제안된 방법의 실용성과 현장 적용 가능성을 입증하였다.

    다양한 환경에서 로봇이 인간의 노동을 효과적으로 대체하기 위해서는, 환경 변화에 따른 성능 안정성이 반드시 보장되어야 한다. 본 논문은 비전 기반 파지의 대표적 난제인 투명성과 클러터 문제를 해결하고, 일반 비전 모듈을 기반으로 한 해법을 제안함으로써 로봇 조작의 안정성과 견고함 향상에 기여하고자 한다.

    더보기

    목차 (Table of Contents)

    • Introduction 5
    • Related Works 14
    • Enhancing Instance Segmentation Modules for Grasping 35
    • Aggregating Multi-View Surface Normals for Grasping 59
    • Interaction-Based Perception 75
    • Introduction 5
    • Related Works 14
    • Enhancing Instance Segmentation Modules for Grasping 35
    • Aggregating Multi-View Surface Normals for Grasping 59
    • Interaction-Based Perception 75
    • Conclusion 83
    • 초록 108
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼