RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    2D 카메라와 적대적 역강화학습(AIRL)을 이용 한 로봇 매니퓰레이터 6 자유도 제어를 위한 마 커리스 인간 모션 모방 = Markerless Human Motion Imitation for Robot Manipulator 6-DoF Control using 2D Camera and Adversarial Inverse Reinforcement Learning

    한글로보기

    https://www.riss.kr/link?id=T17400904

    • 저자
    • 발행사항

      포항 : 한동대학교 일반대학원, 2026

    • 학위논문사항

      학위논문(석사) -- 한동대학교 일반대학원 , 휴먼테크융합학과 , 2026. 2

    • 발행연도

      2026

    • 작성언어

      한국어

    • 발행국(도시)

      경상북도

    • 형태사항

      VI, 31 ; 26 cm

    • 일반주기명

      지도교수: 김재효

    • UCI식별코드

      I804:47030-200000972951

    • 소장기관
      • 한동대학교 도서관 소장기관정보
    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    본 연구에서는 사람의 pick and place 작업을 대상을 해당 작업을 approach, transfer, retreat 의 의미적 행동 모듈(SAM)으로 분해한 후 이를 로봇 매니퓰레이터에서 학습하 고 재현하기 위해 Behavior Cloning(BC)와 Adversarial Inverse Reinforcement Learning(AIRL) 을 결합한 모델을 제안한다. 그리고 제안된 모델을 통해 환경과 인간 행동의 입력에 따라 적절한 SAM 을 이어서 생성할 수 있는 LAM 모델 설계의 기초를 마련한다.
    모델이 인간이 궤적을 선택할 때의 전략과 정책을 학습할 수 있도록 손목과 손끝 의 6 차원 좌표 및 거리 변화율, 이동 방향, 그리고 목표점 근처에서의 이동 방식 등을 보상 함수 설계에 적용했고, 특히 역행과 목표점 근처의 속도를 핵심항으로 사용하여 목표점 도달의 정확도를 높였다. 모델에서 생성된 raw 궤적의 안정적인 적용을 위해, 목표 좌표에 정합 될 수 있도록 재샘플링을 진행하고 grasp 와 placement 작업을 위한 dwell 을 적용하는 후처리를 적용하여 최종적으로 폐곡선 궤적을 완성했다.
    모델 검증은 AIRL-hybrid 와 4 종의 대조군(GAIL-hybrid, BC, MaxEnt)을 동일 환경에 서 학습시킨 후, 산출된 궤적의 RMSE 와 방향 코사인을 통해 기하학적 정합성을 정량 평가하는 방식으로 수행되었다. 분석 결과, AIRL-hybrid 는 대부분의 SAM 구간에서 최저 RMSE 및 최대 방향 코사인을 기록하며 데모 궤적과의 높은 상동성을 입증했으나, Approach 구간에서는 Max Entropy 모델 대비 다소 열세를 보이기도 했다. 그러나 방향 코사인 기반 Confusion Matrix 분석에서 타 모델을 상회하는 최상위 세그먼트 분류 정 확도를 달성한 점은, 제안 모델이 SAM 의 경계를 명확히 식별할 뿐 아니라 내재된 행동 의도까지 정밀하게 모사함을 방증한다. 결과적으로 AIRL-hybrid 는 차세대 로봇 매니퓰레이터의 LAM 구현을 위한 기반 기술로서 충분한 타당성을 확보한 것으로 판단된다.
    번역하기

    본 연구에서는 사람의 pick and place 작업을 대상을 해당 작업을 approach, transfer, retreat 의 의미적 행동 모듈(SAM)으로 분해한 후 이를 로봇 매니퓰레이터에서 학습하 고 재현하기 위해 Behavior Clonin...

    본 연구에서는 사람의 pick and place 작업을 대상을 해당 작업을 approach, transfer, retreat 의 의미적 행동 모듈(SAM)으로 분해한 후 이를 로봇 매니퓰레이터에서 학습하 고 재현하기 위해 Behavior Cloning(BC)와 Adversarial Inverse Reinforcement Learning(AIRL) 을 결합한 모델을 제안한다. 그리고 제안된 모델을 통해 환경과 인간 행동의 입력에 따라 적절한 SAM 을 이어서 생성할 수 있는 LAM 모델 설계의 기초를 마련한다.
    모델이 인간이 궤적을 선택할 때의 전략과 정책을 학습할 수 있도록 손목과 손끝 의 6 차원 좌표 및 거리 변화율, 이동 방향, 그리고 목표점 근처에서의 이동 방식 등을 보상 함수 설계에 적용했고, 특히 역행과 목표점 근처의 속도를 핵심항으로 사용하여 목표점 도달의 정확도를 높였다. 모델에서 생성된 raw 궤적의 안정적인 적용을 위해, 목표 좌표에 정합 될 수 있도록 재샘플링을 진행하고 grasp 와 placement 작업을 위한 dwell 을 적용하는 후처리를 적용하여 최종적으로 폐곡선 궤적을 완성했다.
    모델 검증은 AIRL-hybrid 와 4 종의 대조군(GAIL-hybrid, BC, MaxEnt)을 동일 환경에 서 학습시킨 후, 산출된 궤적의 RMSE 와 방향 코사인을 통해 기하학적 정합성을 정량 평가하는 방식으로 수행되었다. 분석 결과, AIRL-hybrid 는 대부분의 SAM 구간에서 최저 RMSE 및 최대 방향 코사인을 기록하며 데모 궤적과의 높은 상동성을 입증했으나, Approach 구간에서는 Max Entropy 모델 대비 다소 열세를 보이기도 했다. 그러나 방향 코사인 기반 Confusion Matrix 분석에서 타 모델을 상회하는 최상위 세그먼트 분류 정 확도를 달성한 점은, 제안 모델이 SAM 의 경계를 명확히 식별할 뿐 아니라 내재된 행동 의도까지 정밀하게 모사함을 방증한다. 결과적으로 AIRL-hybrid 는 차세대 로봇 매니퓰레이터의 LAM 구현을 위한 기반 기술로서 충분한 타당성을 확보한 것으로 판단된다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This study proposes a hybrid model that combines Behavior Cloning (BC) and Adversarial Inverse Reinforcement Learning (AIRL) to learn and reproduce human pick-and-place motions on a robotic manipulator. Human demonstrations are decomposed into semantic action modules (SAMs)—approach, transfer, and retreat—and the proposed framework is trained to generate these modules sequentially according to the environmental context and observed human behavior, providing a design basis toward future Large Action Model (LAM) implementations.
    Our reward function is primarily grounded in kinematic data derived from demonstrations. By tracking the 6-DoF coordinates of the wrist and fingertips, the system evaluates task progression through distance variance and directional flow. Furthermore, we incorporate shaping terms to strictly penalize backward motion and refine constraints near the target, ensuring high precision in the final positioning.
    Raw rollouts need a cleanup phase to become executable trajectories. We run a post-processing step that resamples data for alignment. It also injects dwell times—pauses—for grasping and releasing. This preserves the original timing structure of the task.
    For validation, AIRL-hybrid went up against GAIL-hybrid, BC, and MaxEnt. All models ran under the exact same training rules. We checked the geometric fit using RMSE and directional cosine. MaxEnt did have a small edge during the approach phase. In most SAM segments, though, AIRLhybrid performed better. The confusion matrix (using directional cosines) backs this up. It shows AIRL-hybrid getting the best classification accuracy, meaning it defines the boundaries between action segments much more sharply than the others
    번역하기

    This study proposes a hybrid model that combines Behavior Cloning (BC) and Adversarial Inverse Reinforcement Learning (AIRL) to learn and reproduce human pick-and-place motions on a robotic manipulator. Human demonstrations are decomposed into semanti...

    This study proposes a hybrid model that combines Behavior Cloning (BC) and Adversarial Inverse Reinforcement Learning (AIRL) to learn and reproduce human pick-and-place motions on a robotic manipulator. Human demonstrations are decomposed into semantic action modules (SAMs)—approach, transfer, and retreat—and the proposed framework is trained to generate these modules sequentially according to the environmental context and observed human behavior, providing a design basis toward future Large Action Model (LAM) implementations.
    Our reward function is primarily grounded in kinematic data derived from demonstrations. By tracking the 6-DoF coordinates of the wrist and fingertips, the system evaluates task progression through distance variance and directional flow. Furthermore, we incorporate shaping terms to strictly penalize backward motion and refine constraints near the target, ensuring high precision in the final positioning.
    Raw rollouts need a cleanup phase to become executable trajectories. We run a post-processing step that resamples data for alignment. It also injects dwell times—pauses—for grasping and releasing. This preserves the original timing structure of the task.
    For validation, AIRL-hybrid went up against GAIL-hybrid, BC, and MaxEnt. All models ran under the exact same training rules. We checked the geometric fit using RMSE and directional cosine. MaxEnt did have a small edge during the approach phase. In most SAM segments, though, AIRLhybrid performed better. The confusion matrix (using directional cosines) backs this up. It shows AIRL-hybrid getting the best classification accuracy, meaning it defines the boundaries between action segments much more sharply than the others

    더보기

    목차 (Table of Contents)

    • Chapter Ⅰ. INTRODUCTION 1
    • Chapter Ⅱ. MATERIALS AND METHODS 4
    • 1. Data Acquisition 4
    • 2. Data Processing and Trajectory Segmentation 5
    • 3. Environment and Data Setup 6
    • Chapter Ⅰ. INTRODUCTION 1
    • Chapter Ⅱ. MATERIALS AND METHODS 4
    • 1. Data Acquisition 4
    • 2. Data Processing and Trajectory Segmentation 5
    • 3. Environment and Data Setup 6
    • 4. Segment-wise BC and AIRL-based Trajectory Learning 11
    • 5. Segment-wise BC and AIRL-based Trajectory Learning 15
    • Chapter III. RESULTS 16
    • 1. Trajectory generation 16
    • 2. Raw rollout properties between BC-only and AIRL-hybrid policies 17
    • 3. Trajectory generation with post-processing 18
    • 4. Generated trajectory analysis 20
    • Chapter IV. DISCUSSION 24
    • 1. Overall Properties of Segment-Based AIRL-Hybrid Policies 24
    • 2. Raw Roll-out Analysis 24
    • 3. Trajectory Generation with Post-Processing 25
    • 4. Imitation Analysis for each Segments 26
    • 5. Limitations of the Study and Future Research Directions 27
    • Chapter V. CONCLUSION 28
    • Chapter Ⅵ. REFERENCE 30
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼