본 연구에서는 고가의 힘/토크 센서나 촉각 센서를 사용하지 않고, RGB 영상과 IMU 기반의 시연 데이터만으로 로봇 조작 행동을 학습하고 실제 UR3 로봇에 재현할 수 있는 센서 최소화 기반 모...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17396770
용인 : 명지대학교 대학원, 2026
2026
한국어
모방학습 ; 멀티모달 트랜스포머 ; 파지 강도 추정 ; 비전 기반 로봇 조작
경기도
A Vision-Based Imitation Learning Framework for Sensorless Robotic Arm Control
vi, 33 p. : 삽화, 도표 ; 26 cm
명지대학교 논문은 저작권에 의해 보호받습니다.
지도교수: 최동일
I804:null-200000949023
0
상세조회0
다운로드본 연구에서는 고가의 힘/토크 센서나 촉각 센서를 사용하지 않고, RGB 영상과 IMU 기반의 시연 데이터만으로 로봇 조작 행동을 학습하고 실제 UR3 로봇에 재현할 수 있는 센서 최소화 기반 모...
본 연구에서는 고가의 힘/토크 센서나 촉각 센서를 사용하지 않고, RGB 영상과 IMU 기반의 시연 데이터만으로 로봇 조작 행동을 학습하고 실제 UR3 로봇에 재현할
수 있는 센서 최소화 기반 모방학습 프레임워크(sensorless imitation learning framework)를 제안한다. 이를 위해 ZED2i 카메라를 장착한 핸드형 그리퍼를 이용하여 RGB 영상, Visual–Inertial SLAM 기반 6-DoF 엔드이펙터 궤적, 그리고 DeepLabCut(DLC)을 활용한 그리퍼 팁 변형률을 수집하였으며, 이러한 정보를 기반으
로 새로운 비디오·궤적·파지 강도 데이터셋을 구축하였다. 제안한 모델은 트랜스포머 기반 멀티모달 구조로, 비전 특성(ViT), 궤적 시계열(GRU), 파지 상태 정보(grasp action, grasp strength), 작업명(task label)을 통합하여 6-DoF 궤적과 연속 파지 강도를 동시에 예측하는 멀티태스크 학습 모델로 설계되었다. 학습된
모델은 궤적 예측에서 3D RMSE 12.40 mm, 파지 강도 추정에서 MAE 0.0111, RMSE 0.0155의 성능을 기록하여 실제 조작 작업에 적합한 수준의 정확도를 달성하였다. 실제 UR3 로봇을 이용한 검증 실험에서는 생성된 궤적을 역기구학(IK)을 통해 관절각으로 변환하여 open-loop 방식으로 실행하였으며, Pick & Place 작업에서 95% 성공률(20회 중 19회)을 달성하였다. 또한 물체별 파지 강도 실험에서도 네 종류의 물체에 대해 최대 100%의 성공률을 보이며, 영상 기반 변형률 추정만으로도 안정적인 센서리스 파지 제어가 가능함을 확인하였다. 본 연구는 단일 카메라와 IMU만으로 로봇 조작 데이터를 획득하고 로봇 작업을 재현할 수 있음을 실험적으로 증명하였으며, 저비용·저복잡도 기반의 범용 모방학습 로봇 제어 시스템 구축 가능성을 제시한다.
다국어 초록 (Multilingual Abstract)
In this study, we propose a sensor-minimal imitation learning framework that enables robot manipulation using only RGB video and IMU-based demonstration data, without relying on force or tactile sensors. A handheld gripper with a ZED2i camera was use...
In this study, we propose a sensor-minimal imitation learning framework that enables robot manipulation using only RGB video and IMU-based demonstration data, without relying on force or tactile sensors. A handheld gripper with a ZED2i
camera was used to collect RGB video, Visual–Inertial SLAM–based 6-DoF trajectories, and gripper tip deformation extracted via DeepLabCut, forming a multimodal dataset of video, trajectories, and grasp-strength labels. The proposed transformer-based multimodal model integrates visual features (ViT), trajectory signals (GRU), grasp states, and task labels to jointly predict 6-DoF poses and continuous grasp strength. The model achieves 12.40 mm trajectory RMSE and 0.0111 MAE for grasp strength, demonstrating accuracy suitable for manipulation tasks. For real-world validation, predicted trajectories were converted into joint angles via inverse kinematics and executed on a UR3 robot, achieving a 95% success rate in pick-and-place tasks and up to 100% success in object-wise grasp tests. These results show that effective robot manipulation can be learned and reproduced using only a single RGB camera and IMU, highlighting the practicality of low-cost, sensorless robot control.
목차 (Table of Contents)