RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Integrating Flow Matching into Decision Transformers for Enhanced Offline Reinforcement Learning = 오프라인 강화학습 성능 향상을 위한 의사결정 트랜스포머와 플로우 매칭의 통합

    한글로보기

    https://www.riss.kr/link?id=T17503632

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    The Decision Transformer has brought innovation to the field of offline reinforcement learning by reapproaching the trajectory optimization problem from a sequence modeling perspective. This approach enhances system scalability, enables effective credit assignment over long-term dependencies, and demonstrates excellent generalization performance across various domains. However, it faces two significant limitations. First, sample-based learning, which relies on offline data, covers only a limited portion of the state-action space, leading to compounding errors when encountering out-of-distribution situations. Second, its autoregressive architecture exacerbates these errors, leading to inaccuracies that accumulate and reduce robustness in long-horizon tasks.

    Meanwhile, Flow Matching has emerged as a powerful generative modeling technique. It learns to transfer probability mass from a simple prior distribution to a complex data distribution through a continuous vector field. This process bypasses the need for stochastic sampling and likelihood estimation, enabling faster inference and more stable training while maintaining expressive representational power.

    Building on these insights, this paper proposes a novel integrated framework, DT+FM, which combines the strengths of Decision Transformer and Flow Matching. By replacing sample-based learning with distribution-based learning, DT+FM broadens coverage of the state-action space and reduces compounding errors. Experimental results in continuous control environments using MuJoCo and Robomimic show that the proposed DT+FM consistently outperforms the original Decision Transformer, particularly demonstrating greater stability and consistent performance across tasks characterized by long episodes and complex dynamics. These results highlight DT+FM’s promise as a compelling direction for advancing offline reinforcement learning.
    번역하기

    The Decision Transformer has brought innovation to the field of offline reinforcement learning by reapproaching the trajectory optimization problem from a sequence modeling perspective. This approach enhances system scalability, enables effective cred...

    The Decision Transformer has brought innovation to the field of offline reinforcement learning by reapproaching the trajectory optimization problem from a sequence modeling perspective. This approach enhances system scalability, enables effective credit assignment over long-term dependencies, and demonstrates excellent generalization performance across various domains. However, it faces two significant limitations. First, sample-based learning, which relies on offline data, covers only a limited portion of the state-action space, leading to compounding errors when encountering out-of-distribution situations. Second, its autoregressive architecture exacerbates these errors, leading to inaccuracies that accumulate and reduce robustness in long-horizon tasks.

    Meanwhile, Flow Matching has emerged as a powerful generative modeling technique. It learns to transfer probability mass from a simple prior distribution to a complex data distribution through a continuous vector field. This process bypasses the need for stochastic sampling and likelihood estimation, enabling faster inference and more stable training while maintaining expressive representational power.

    Building on these insights, this paper proposes a novel integrated framework, DT+FM, which combines the strengths of Decision Transformer and Flow Matching. By replacing sample-based learning with distribution-based learning, DT+FM broadens coverage of the state-action space and reduces compounding errors. Experimental results in continuous control environments using MuJoCo and Robomimic show that the proposed DT+FM consistently outperforms the original Decision Transformer, particularly demonstrating greater stability and consistent performance across tasks characterized by long episodes and complex dynamics. These results highlight DT+FM’s promise as a compelling direction for advancing offline reinforcement learning.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    Decision Transformer는 궤적 최적화(trajectory optimizartion) 문제를 시퀀스 모델링의 관점에서 새롭게 접근함으로써 오프라인 강화학습 분야에 혁신을 가져왔다. 이는 시스템의 확장성을 높이고, 장기 의존성에 대한 효과적인 신뢰 할당을 가능하게 하며, 다양한 도메인에서 우수한 일반화 성능을 나타낸다. 그러나 이 접근 방식에는 두 가지 중요한 한계점이 존재한다. 첫째, 오프라인 데이터에 기반한 샘플 중심 학습은 상태–행동 공간의 제한적인 부분만을 학습하게 되어, 학습 분포에서 벗어난 상황에서 오류가 누적되는(compounding error) 문제가 발생한다. 둘째, 자기회귀적(autoregressive) 아키텍처는 이러한 오류를 한층 더 증폭시켜 부정확성이 반복적으로 쌓이게 되며, 그 결과 장기 과제에서의 강건성이 감소한다.

    한편, Flow Matching은 최근 주목받고 있는 강력한 생성 모델링 기법이다. Flow Matching은 단순한 사전 분포로부터 복잡한 데이터 분포로 확률 질량을 연속 벡터장으로 전달하는 방식을 학습한다. 이 과정은 확률적 샘플링과 우도 추정 단계를 거치지 않게 하여, 더 빠른 추론과 안정적인 학습이 가능하게 한다. 또한, 이 방법은 강력한 표현력을 유지한다. 이러한 통찰에 기반하여, 본 논문은 Decision Transformer와 Flow Matching의 장점을 통합한 DT+FM이라는 새로운 통합 프레임워크를 제안한다. 샘플 기반 학습을 분포 기반 학습으로 대체함으로써 DT+FM은 상태–행동 공간의 범위를 넓히고 누적 오류를 줄이는 데 기여한다. MuJoCo와 Robomimic 기반 연속 제어 환경에서 수행한 실험 결과, 제안된 DT+FM은 기존 Decision Transformer보다 일관되게 우수한 성능을 보였으며, 특히 에피소드 길이가 길고 복잡한 동역학을 가지는 과제에서 더 높은 안정성과 일관된 성능을 나타냈다. 이러한 결과는 오프라인 강화학습 발전을 위한 유망한 방향으로서 DT+FM의 가능성을 보여준다.
    번역하기

    Decision Transformer는 궤적 최적화(trajectory optimizartion) 문제를 시퀀스 모델링의 관점에서 새롭게 접근함으로써 오프라인 강화학습 분야에 혁신을 가져왔다. 이는 시스템의 확장성을 높이고, 장...

    Decision Transformer는 궤적 최적화(trajectory optimizartion) 문제를 시퀀스 모델링의 관점에서 새롭게 접근함으로써 오프라인 강화학습 분야에 혁신을 가져왔다. 이는 시스템의 확장성을 높이고, 장기 의존성에 대한 효과적인 신뢰 할당을 가능하게 하며, 다양한 도메인에서 우수한 일반화 성능을 나타낸다. 그러나 이 접근 방식에는 두 가지 중요한 한계점이 존재한다. 첫째, 오프라인 데이터에 기반한 샘플 중심 학습은 상태–행동 공간의 제한적인 부분만을 학습하게 되어, 학습 분포에서 벗어난 상황에서 오류가 누적되는(compounding error) 문제가 발생한다. 둘째, 자기회귀적(autoregressive) 아키텍처는 이러한 오류를 한층 더 증폭시켜 부정확성이 반복적으로 쌓이게 되며, 그 결과 장기 과제에서의 강건성이 감소한다.

    한편, Flow Matching은 최근 주목받고 있는 강력한 생성 모델링 기법이다. Flow Matching은 단순한 사전 분포로부터 복잡한 데이터 분포로 확률 질량을 연속 벡터장으로 전달하는 방식을 학습한다. 이 과정은 확률적 샘플링과 우도 추정 단계를 거치지 않게 하여, 더 빠른 추론과 안정적인 학습이 가능하게 한다. 또한, 이 방법은 강력한 표현력을 유지한다. 이러한 통찰에 기반하여, 본 논문은 Decision Transformer와 Flow Matching의 장점을 통합한 DT+FM이라는 새로운 통합 프레임워크를 제안한다. 샘플 기반 학습을 분포 기반 학습으로 대체함으로써 DT+FM은 상태–행동 공간의 범위를 넓히고 누적 오류를 줄이는 데 기여한다. MuJoCo와 Robomimic 기반 연속 제어 환경에서 수행한 실험 결과, 제안된 DT+FM은 기존 Decision Transformer보다 일관되게 우수한 성능을 보였으며, 특히 에피소드 길이가 길고 복잡한 동역학을 가지는 과제에서 더 높은 안정성과 일관된 성능을 나타냈다. 이러한 결과는 오프라인 강화학습 발전을 위한 유망한 방향으로서 DT+FM의 가능성을 보여준다.

    더보기

    목차 (Table of Contents)

    • 1. Introduction 1
    • 1.1. Background and Motivation 1
    • 1.2. Problem Statement 2
    • 1.3. Research Contributions 3
    • 1.4. Thesis Structure 3
    • 1. Introduction 1
    • 1.1. Background and Motivation 1
    • 1.2. Problem Statement 2
    • 1.3. Research Contributions 3
    • 1.4. Thesis Structure 3
    • 2. Related Works 5
    • 2.1. Offline Reinforcement Learning from Fixed Datasets 5
    • 2.2. Decision Transformer and Its Extensions 7
    • 2.3. Addressing Decision Transformer Limitations with Generative Models 9
    • 2.4. Flow Matching in Reinforcement Learning 13
    • 3. Preliminaries 15
    • 3.1. Offline Reinforcement Learning 15
    • 3.2. Decision Transformer and Sequence Modeling for Control 16
    • 3.3. Flow Matching 19
    • 3.3.1. Flow ODE and Pushforward Density 19
    • 3.3.2. Conditional Probability Paths and the FM Objective 20
    • 3.3.3. Conditional FM and Comparison to Diffusion 21
    • 4. Decision Transformer with Integrated Flow Matching (DT+FM) 22
    • 4.1. Addressing Compounding Error and Distribution Shift via Flow Matching Integration 22
    • 4.2. DT+FM Architecture 23
    • 4.3. Training and Inference Pipelines 26
    • 5. Experiments 28
    • 5.1. Environments and Benchmarks 28
    • 5.1.1. MuJoCo Tasks 28
    • 5.1.2. RoboMimic Tasks 31
    • 5.2. Baseline for Comparison 32
    • 5.3. Experimental Setup 33
    • 5.4. Experiment Results 35
    • 5.4.1. MuJoCo Locomotion Results 35
    • 5.4.2. RoboMimic Tasks Results 40
    • 6. Conclusion 42
    • Bibliography 44
    • Abstract 50
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼