RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Generative Modeling of 3D Humans with Geometry-Aware Diffusion Models = 기하 정보를 고려한 확산 모델 기반 3D 인간의 생성적 모델링

    한글로보기

    https://www.riss.kr/link?id=T17450372

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    3D human modeling aims to reconstruct 3D human geometry from limited visual observations and to generate images or videos that are consistent with underlying 3D structure. Despite substantial progress, reliable human-centric 3D modeling remains challenging due to the scarcity of high-quality 3D and 4D supervision, the articulated and non-rigid nature of the human body, and the increasing ambiguity encountered in stylized domains, multi-human scenes, and dynamic settings. While large-scale 2D image and video datasets are widely available, directly inferring or enforcing consistent 3D structure from such data remains fundamentally underconstrained without additional structural guidance.

    This thesis studies geometry-aware diffusion models as a practical framework for human-centric 3D modeling under limited supervision. The central idea is to combine diffusion-based generative models, which learn strong appearance, structure, and motion representations from large-scale 2D data, with feasible geometric guidance such as pose, depth, surface normals, and visibility cues. These geometric signals serve to bridge 2D diffusion models and 3D structure, reducing ambiguity during generation and reconstruction.

    The thesis is organized around three progressively more challenging scenarios. Part I addresses stylized and novel-domain 3D human and character generation, where target-domain geometric annotations are unavailable. It introduces diffusion-based pose-aware data generation and pose-preserved adaptation strategies that enable pretrained 3D generative models to adapt to new appearance domains while maintaining geometric consistency. Part II focuses on multi-human scenes, where occlusion, depth ordering ambiguity, and inter-person interaction complicate both generation and reconstruction. This part presents interaction-aware diffusion frameworks that incorporate occlusion-aware and group-level geometric representations to support coherent multi-human image synthesis and 3D reconstruction from limited visual input. Part III considers dynamic human scenes and studies temporally consistent 3D geometry estimation from monocular video. By reformulating dynamic geometry inference as a geometry-aware image-to-video diffusion process, this part leverages temporal priors learned by video diffusion models to stabilize geometry over time, reducing reliance on costly 4D supervision.

    Overall, this thesis demonstrates that geometry-aware diffusion models provides a practical and extensible approach for human-centric 3D modeling under data scarcity. By selecting, controlling, and structuring geometric guidance according to the dominant source of uncertainty in each setting, the proposed framework enables coherent, controllable, and temporally stable 3D human generation and reconstruction across a wide range of challenging scenarios.
    번역하기

    3D human modeling aims to reconstruct 3D human geometry from limited visual observations and to generate images or videos that are consistent with underlying 3D structure. Despite substantial progress, reliable human-centric 3D modeling remains challe...

    3D human modeling aims to reconstruct 3D human geometry from limited visual observations and to generate images or videos that are consistent with underlying 3D structure. Despite substantial progress, reliable human-centric 3D modeling remains challenging due to the scarcity of high-quality 3D and 4D supervision, the articulated and non-rigid nature of the human body, and the increasing ambiguity encountered in stylized domains, multi-human scenes, and dynamic settings. While large-scale 2D image and video datasets are widely available, directly inferring or enforcing consistent 3D structure from such data remains fundamentally underconstrained without additional structural guidance.

    This thesis studies geometry-aware diffusion models as a practical framework for human-centric 3D modeling under limited supervision. The central idea is to combine diffusion-based generative models, which learn strong appearance, structure, and motion representations from large-scale 2D data, with feasible geometric guidance such as pose, depth, surface normals, and visibility cues. These geometric signals serve to bridge 2D diffusion models and 3D structure, reducing ambiguity during generation and reconstruction.

    The thesis is organized around three progressively more challenging scenarios. Part I addresses stylized and novel-domain 3D human and character generation, where target-domain geometric annotations are unavailable. It introduces diffusion-based pose-aware data generation and pose-preserved adaptation strategies that enable pretrained 3D generative models to adapt to new appearance domains while maintaining geometric consistency. Part II focuses on multi-human scenes, where occlusion, depth ordering ambiguity, and inter-person interaction complicate both generation and reconstruction. This part presents interaction-aware diffusion frameworks that incorporate occlusion-aware and group-level geometric representations to support coherent multi-human image synthesis and 3D reconstruction from limited visual input. Part III considers dynamic human scenes and studies temporally consistent 3D geometry estimation from monocular video. By reformulating dynamic geometry inference as a geometry-aware image-to-video diffusion process, this part leverages temporal priors learned by video diffusion models to stabilize geometry over time, reducing reliance on costly 4D supervision.

    Overall, this thesis demonstrates that geometry-aware diffusion models provides a practical and extensible approach for human-centric 3D modeling under data scarcity. By selecting, controlling, and structuring geometric guidance according to the dominant source of uncertainty in each setting, the proposed framework enables coherent, controllable, and temporally stable 3D human generation and reconstruction across a wide range of challenging scenarios.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    본 논문은 기하 정보를 고려한 확산 모델 기반 3D 인간의 생성적 모델링을 주제로, 제한된 3D 및 4D 감독 환경에서의 인간 중심 3D 모델링 문제를 다룬다. 3D 인간 모델링은 단일 시점 이미지나 비디오와 같은 제한된 시각적 관측으로부터 3차원 인간 기하를 복원하거나, 기저의 3D 구조와 일관된 이미지 및 비디오를 생성하는 것을 목표로 한다. 그러나 고품질 3D·4D 데이터의 희소성, 인간 신체의 관절 구조와 비강체 특성, 그리고 스타일화된 도메인, 다중 인물 장면, 동적 환경에서 증가하는 구조적 모호성으로 인해 신뢰성 있는 3D 인간 모델링은 여전히 어려운 문제로 남아 있다. 대규모 2D 이미지 및 비디오 데이터는 풍부하게 존재하지만, 추가적인 구조적 제약 없이 이러한 데이터로부터 일관된 3D 구조를 추론하거나 강제하는 것은 본질적으로 도전적인 문제이다.

    본 논문은 이러한 문제를 해결하기 위한 통합적 접근으로 기하 정보를 고려한 확산 모델을 제안한다. 핵심 아이디어는 대규모 2D 데이터로부터 외형, 구조, 동작, 의미적 표현을 학습한 확산 기반 생성 모델에 자세, 깊이, 표면 법선, 가시성 정보와 같은 실현 가능한 기하적 가이드를 결합하는 것이다. 이러한 기하 정보는 2D 확산 모델과 3D 구조 사이를 연결하는 부분적 제약으로 작용하여 생성 및 복원 과정에서의 모호성을 효과적으로 줄인다.

    본 논문은 난이도가 점진적으로 증가하는 세 가지 시나리오를 중심으로 구성된다. 첫째, Part I에서는 타깃 도메인에 기하 주석이 존재하지 않는 스타일화된 및 신규 도메인의 3D 인간 및 캐릭터 생성 문제를 다룬다. 자세 인지 데이터 생성과 자세 보존 확산 모델을 통해, 사전 학습된 3D 생성 모델이 기하적 일관성을 유지한 채 새로운 외형 도메인으로 적응할 수 있음을 보인다. 둘째, Part II에서는 가림 현상, 깊이 순서 모호성, 인물 간 상호작용으로 인해 생성과 복원이 복잡해지는 다중 인물 장면을 다룬다. 본 부분에서는 가림을 고려한 기하 표현과 그룹 수준의 구조를 통합한 상호작용 인지 확산 프레임워크를 제시하여, 제한된 시각 입력만으로도 일관된 다중 인물 이미지 생성과 3D 장면 복원을 가능하게 한다. 셋째, Part III에서는 단안 비디오로부터 시간적으로 일관된 3D 기하 추정을 목표로 하는 동적 인간 장면을 다룬다. 동적 기하 추정을 기하 정보를 고려한 이미지-비디오 확산 문제로 재정식화함으로써, 비디오 확산 모델이 학습한 시간적 사전 지식을 활용하여 4D 감독에 대한 의존도를 줄이면서 안정적인 기하 추정을 달성한다.

    종합적으로 본 논문은 기하 정보를 고려한 확산 모델이 데이터가 제한된 환경에서도 3D 인간의 생성과 복원을 효과적으로 통합할 수 있는 실용적이고 확장 가능한 접근임을 보인다. 각 문제 설정에서 지배적인 불확실성의 원인에 맞추어 기하 정보를 선택하고, 조절하며, 구조화함으로써, 제안된 프레임워크는 다양한 도전적 환경에서 일관성 있고 제어 가능하며 시간적으로 안정적인 3D 인간 모델링을 가능하게 한다.
    번역하기

    본 논문은 기하 정보를 고려한 확산 모델 기반 3D 인간의 생성적 모델링을 주제로, 제한된 3D 및 4D 감독 환경에서의 인간 중심 3D 모델링 문제를 다룬다. 3D 인간 모델링은 단일 시점 이미지나 ...

    본 논문은 기하 정보를 고려한 확산 모델 기반 3D 인간의 생성적 모델링을 주제로, 제한된 3D 및 4D 감독 환경에서의 인간 중심 3D 모델링 문제를 다룬다. 3D 인간 모델링은 단일 시점 이미지나 비디오와 같은 제한된 시각적 관측으로부터 3차원 인간 기하를 복원하거나, 기저의 3D 구조와 일관된 이미지 및 비디오를 생성하는 것을 목표로 한다. 그러나 고품질 3D·4D 데이터의 희소성, 인간 신체의 관절 구조와 비강체 특성, 그리고 스타일화된 도메인, 다중 인물 장면, 동적 환경에서 증가하는 구조적 모호성으로 인해 신뢰성 있는 3D 인간 모델링은 여전히 어려운 문제로 남아 있다. 대규모 2D 이미지 및 비디오 데이터는 풍부하게 존재하지만, 추가적인 구조적 제약 없이 이러한 데이터로부터 일관된 3D 구조를 추론하거나 강제하는 것은 본질적으로 도전적인 문제이다.

    본 논문은 이러한 문제를 해결하기 위한 통합적 접근으로 기하 정보를 고려한 확산 모델을 제안한다. 핵심 아이디어는 대규모 2D 데이터로부터 외형, 구조, 동작, 의미적 표현을 학습한 확산 기반 생성 모델에 자세, 깊이, 표면 법선, 가시성 정보와 같은 실현 가능한 기하적 가이드를 결합하는 것이다. 이러한 기하 정보는 2D 확산 모델과 3D 구조 사이를 연결하는 부분적 제약으로 작용하여 생성 및 복원 과정에서의 모호성을 효과적으로 줄인다.

    본 논문은 난이도가 점진적으로 증가하는 세 가지 시나리오를 중심으로 구성된다. 첫째, Part I에서는 타깃 도메인에 기하 주석이 존재하지 않는 스타일화된 및 신규 도메인의 3D 인간 및 캐릭터 생성 문제를 다룬다. 자세 인지 데이터 생성과 자세 보존 확산 모델을 통해, 사전 학습된 3D 생성 모델이 기하적 일관성을 유지한 채 새로운 외형 도메인으로 적응할 수 있음을 보인다. 둘째, Part II에서는 가림 현상, 깊이 순서 모호성, 인물 간 상호작용으로 인해 생성과 복원이 복잡해지는 다중 인물 장면을 다룬다. 본 부분에서는 가림을 고려한 기하 표현과 그룹 수준의 구조를 통합한 상호작용 인지 확산 프레임워크를 제시하여, 제한된 시각 입력만으로도 일관된 다중 인물 이미지 생성과 3D 장면 복원을 가능하게 한다. 셋째, Part III에서는 단안 비디오로부터 시간적으로 일관된 3D 기하 추정을 목표로 하는 동적 인간 장면을 다룬다. 동적 기하 추정을 기하 정보를 고려한 이미지-비디오 확산 문제로 재정식화함으로써, 비디오 확산 모델이 학습한 시간적 사전 지식을 활용하여 4D 감독에 대한 의존도를 줄이면서 안정적인 기하 추정을 달성한다.

    종합적으로 본 논문은 기하 정보를 고려한 확산 모델이 데이터가 제한된 환경에서도 3D 인간의 생성과 복원을 효과적으로 통합할 수 있는 실용적이고 확장 가능한 접근임을 보인다. 각 문제 설정에서 지배적인 불확실성의 원인에 맞추어 기하 정보를 선택하고, 조절하며, 구조화함으로써, 제안된 프레임워크는 다양한 도전적 환경에서 일관성 있고 제어 가능하며 시간적으로 안정적인 3D 인간 모델링을 가능하게 한다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents iii
    • List of Figures vi
    • Abstract i
    • Contents iii
    • List of Figures vi
    • List of Tables xi
    • 1 Introduction 1
    • 1.1 Challenges in Generative Modeling of 3D Humans 2
    • 1.2 Thesis Statement and Contributions 5
    • 1.3 Thesis Organization 7
    • I PART I: 3D HUMAN AND CHARACTER GENERATION 10
    • 2 Pose-Aware Synthetic Data Generation for Domain Adaptation 11
    • 2.1 Problem Formulation 12
    • 2.2 Methodology 15
    • 2.3 Experimental Results 20
    • 3 Pose-Preserving Diffusion Models for Novel Domains 25
    • 3.1 Architecture Overview 26
    • 3.2 Loss Functions and Training 31
    • 3.3 Evaluation 36
    • II PART II: ROBUST 3D RECONSTRUCTION AND STYLE TRANSFER 42
    • 4 Geometry-Guided Feature Fusion for Appearance Transfer 43
    • 4.1 Attention-Based Feature Injection 45
    • 4.2 Preservation of Semantic Identity 52
    • 5 3D Reconstruction from Sparse Views with Structural Priors 60
    • 5.1 Joint Optimization of Depth and Normal 62
    • 5.2 Refining Detailed Surface Geometry 71
    • III PART III: COMPOSITIONAL AND DYNAMIC HUMAN GENERATION 85
    • 6 Diffusion Models for Multi-Human Scene Synthesis 86
    • 6.1 Spatial Control via Bounding Boxes 89
    • 6.2 Disentangling Occlusion and Visibility 101
    • 7 Compositional Video Generation with Consistent Dynamics 115
    • 7.1 Temporal Coherence in World Models 118
    • 7.2 Quantitative and Qualitative Comparisons 124
    • 8 Conclusion 140
    • 8.1 Summary of Contributions 141
    • 8.2 Future Research Directions 143
    • Bibliography 148
    • Abstract (Korean) 165
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼