강화학습은 로봇팔 조작부터 보행까지 다양한 작업에서 로봇 제어의 핵심 패러다임으로 자리 잡고 있다. 최근 발전은 강화학습이 장애물 회피나 비파지 조작과 같은 복잡한 문제도 어려운 ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
강화학습은 로봇팔 조작부터 보행까지 다양한 작업에서 로봇 제어의 핵심 패러다임으로 자리 잡고 있다. 최근 발전은 강화학습이 장애물 회피나 비파지 조작과 같은 복잡한 문제도 어려운 ...
강화학습은 로봇팔 조작부터 보행까지 다양한 작업에서 로봇 제어의 핵심 패러다임으로 자리 잡고 있다. 최근 발전은 강화학습이 장애물 회피나 비파지 조작과 같은 복잡한 문제도 어려운 환경에서 해결할 수 있음을 보여준다. 그러나 일반적인 강화학습으로 학습된 정책은 훈련 환경과 테스트 환경 사이에 차이가 발생하면 성능이 저하되는 문제가 있다. 이를 완화하기 위해 학습 단계에서 다양한 환경에 노출시켜 여러 환경에서의 일반화를 도모하는 도메인 무작위화가 널리 사용된다. 하지만 이 방법은 무작위화된 훈련 분포 내부에서는 강건성을 보이더라도, 분포 밖 새로운 환경에서 적절한 행동을 보장하지 못한다. 예컨대 조작 작업에서는 물체의 질량·마찰처럼 시각적으로 관찰할 수 없는 잠재 물리 특성들을 무작위화하고, 동작 계획에서는 다양한 장애물 구성을 통해 정책이 충돌을 피하는 동작을 학습하게 한다. 그럼에도 불구하고 훈련 분포 내 강건함과 달리, 정책이 물체의 실제 물리 특성에 맞는 행동을 선택하거나, 새로운/동적으로 변화하는 장애물에 대해 안전한 동작 계획을 수행할 것이라는 보장은 없다. 이 한계는 로봇이 불확실한 물체와 환경에서 다양한 행동을 선택·적응할 수 있도록 하는 접근의 필요성을 강조한다.
본 논문은 이러한 한계를 극복하기 위한 정책 학습 방법을 제시한다. 우리의 접근법은 훈련 중 정책을 다양한 환경에 노출하여 과제를 성공으로 이끄는 여러 행동을 학습하게 하고, 그 행동들을 잠재적 기술 (latent skill) 로 인코딩한다. 정책은 상황에 따라 적절한 잠재 기술을 선택함으로써 하나의 정책으로도 여러 가지 행동을 수행할 수 있다. 환경 다양화 자체는 전통적 강화학습과 유사하나, 본 방법은 정책이 명시적인 행동 레퍼토리를 갖추도록 설계된다는 점에서 차별화된다. 그 결과, 불확실한 물체와 환경에 대해 단일 견고 정책을 학습하는 방식보다 우수한 성능을 보인다.
이 개념을 바탕으로 우리는 장애물 회피를 위한 새로운 강화학습 기반 동작 계획 방법을 제안한다. 훈련 중 장애물을 다양한 위치에 무작위 배치하여 가능한 행동 공간을 의도적으로 제약함으로써, 정책이 단일 전략에 의존하지 않고 여러 회피 행동을 탐색·습득하도록 유도한다. 학습된 다양한 해결책은 잠재 기술로 인코딩되며, 테스트 시에는 두 가지 핵심 요소-잠재 기술 샘플러와 모션 예측기-가 추가로 작동한다. 잠재 기술 샘플러가 후보 기술을 생성하면, 모션 예측기가 결과 궤적을 예측해 안전하지 않거나 비효율적인 후보를 걸러낸다. 이 메커니즘을 통해 추가 학습 없이도 새로운 장애물이나 동적으로 변화하는 장애물에 제로샷 적응이 가능하다. 실험 결과, 제안 방법은 동적 제약 환경에서 기존 방법 대비 높은 적응성과 성능을 보임을 확인했다.
다음으로 우리는 숨겨진 물리적 특성의 변동성 하에서 물체 재배치를 위한 다양한 조작 기술 (파지 및 비파지 조작) 을 학습하는 방법을 제안한다. 이를 위해 무작위 물체 포인트 클라우드와 그리퍼 자세를 입력으로 접촉과 파지 가능성을 예측하도록 학습된 파지 가능성 인식 접촉 표현을 도입한다. 이 표현은 접촉·파지 가능성 정보를 효과적으로 인코딩하며, 훈련 중에는 질량, 마찰, 질량 중심 등 잠재 물리 매개변수를 무작위화해 변동성을 유도하고, 정책이 파지 및 비파지 조작을 포괄하는 다양한 기술을 습득하도록 한다. 테스트 시에는 실제 물리 특성이 관측되지 않으므로, 잠재 기술 추론 모듈을 통해 로봇·물체 상태와 짧은 행동 이력으로 잠재 기술을 추정한다. 이를 바탕으로 기하학적 구조 변화에 맞춰 기술을 실시간 적응·전환할 수 있게 한다.
끝으로, 시뮬레이션에서 학습한 정책을 실제 환경에 적용할 때 발생하는 시뮬레이션–현실 분포 이동 문제를 해결하기 위해 물리 기반 시뮬레이터 비의존적 보정 방법을 제안한다. 실제 환경과 시뮬레이션 환경에서 반복 가능한 로봇팔 탐색 동작을 수행해 데이터를 수집하고, 시뮬레이터 매개변수를 최적화한다. 구체적으로 경로 및 접촉 기술자 (descriptor) 를 도입해 양쪽 결과를 정량 비교한 뒤, 그리드 탐색으로 최적 매개변수를 찾는다. 이 보정 절차는 탐색 과정에서 생성된 물체 경로와 물리적 접촉의 차이를 효과적으로 줄여 주며, 보정된 설정은 시뮬레이션된 운동학을 현실 세계와 더 정확히 정렬한다. 결과적으로, 학습된 정책의 실제 배포 신뢰성이 크게 향상됨을 확인했다.
다국어 초록 (Multilingual Abstract)
Reinforcement learning (RL) has emerged as a central paradigm for robot control across tasks from robotic manipulation to legged locomotion. Recent progress has shown that RL can tackle complex problems-such as obstacle avoidance and non-prehensile ma...
Reinforcement learning (RL) has emerged as a central paradigm for robot control across tasks from robotic manipulation to legged locomotion. Recent progress has shown that RL can tackle complex problems-such as obstacle avoidance and non-prehensile manipulation-even in challenging environments. However, policies trained with conventional RL often exhibit performance degradation when there is a mismatch between the training and test environments (i.e., under train–test distribution shift). To mitigate this, domain randomization, which exposes a policy to diverse environments during training to promote generalization across settings is widely adopted. Yet, while DR can yield robustness within the randomized training distribution, it does not guarantee appropriate behavior in out-of-distribution environments. For example, in manipulation tasks, latent physical properties-such as mass and friction-that are not directly observable from visual observations are randomized; in motion planning, diverse obstacle configurations are used to encourage the policy to learn collision-free motions. Nevertheless, robustness within the training distribution does not ensure that the policy will select actions consistent with an object’s true physical properties, nor does it guarantee safe motion plans under novel or dynamically changing obstacles. This limitation underscores the need for approaches that enable robots to select among diverse behaviors under uncertainty in both objects and environments.
In this thesis, we introduce a policy-learning framework designed to overcome these limitations. Our approach exposes the policy to a variety of environments during training, thereby inducing the acquisition of multiple task-solving behaviors that are then encoded as latent skills. At deployment, a single policy executes diverse behaviors by selecting an appropriate latent skill conditioned on the current context. While such environmental variation resembles conventional RL training, our method explicitly equips the policy with a behavior repertoire, which differentiates our method from standard pipelines. Consequently, it outperforms training a single “robust” policy intended to generalize across uncertain object properties and environmental conditions.
Building on this idea, we propose a new reinforcement learning–based motion planning method for obstacle avoidance. During training, obstacles are randomly instantiated at varying locations to deliberately constrain the feasible action space, inducing exploration and acquisition of multiple avoidance behaviors rather than reliance on a single strategy. The learned diverse solutions are encoded as latent skills, and at test time we additionally employ two core components-a latent skill sampler and a motion predictor. The latent skill sampler proposes candidate skills, and the motion predictor forecasts the resulting trajectories to filter out candidates that are unsafe or inefficient. This mechanism enables zero-shot adaptation to novel or time-varying obstacles without additional training. In experiments, the proposed method outperforms prior approaches, demonstrating superior adaptability and performance in environments with dynamic constraints.
Next, we propose a method for learning a repertoire of manipulation skills-including grasping and non-prehensile actions-for object rearrangement under variability in latent physical properties. To this end, we introduce a graspability-aware contact representation trained by predicting contact and graspability from randomized object point clouds and gripper poses. This representation effectively encodes contact and graspability information. During training, analogous to the previous method, we randomize latent physical parameters-mass, friction, and center of mass-to induce variability and drive acquisition of a diverse skill set covering both grasping and non-prehensile manipulation. At test time, since the ground-truth physical properties are unobserved, we incorporate a latent skill inference module that estimates the latent variables from short histories of robot states, object states, and actions. Based on these estimates, the policy performs real-time adaptation and skill switching in response to changes in object poses.
Finally, to address the simulation-to-reality (sim-to-real) distribution shift that arises when deploying policies trained in simulation, we propose a physics-oriented, simulator-agnostic calibration method. We execute repeatable probing motions with a robotic manipulator in both the real world and simulation to collect paired data, and we optimize simulator parameters accordingly. Specifically, we introduce trajectory- and contact-level descriptors to quantitatively compare outcomes between the two domains, and we perform a grid search to identify optimal parameter settings. This calibration procedure reduces discrepancies in object trajectories and physical contacts observed during the probes, and the calibrated configuration aligns simulated kinematics more closely with kinematics measured in the real world. Consequently, we observe a substantial improvement in the reliability of real-world deployment of the learned policies.
목차 (Table of Contents)