본 학위논문은 점 군 데이터를 서로 다른 문제 설정에서 활용하는 두 가지 연구 방향을 제시한다. 점 군은 컴퓨터 비전 및 관련 분야에서 기하학적 구조를 표현하기 위한 널리 사용되는 데이...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17450003
서울 : 서울대학교 대학원, 2026
학위논문(박사) -- 서울대학교 대학원 , 협동과정인공지능전공 , 2026. 2
2026
영어
006.3
서울
ix, 121 ; 26 cm
지도교수: 강명주
I804:11032-000000193457
0
상세조회0
다운로드본 학위논문은 점 군 데이터를 서로 다른 문제 설정에서 활용하는 두 가지 연구 방향을 제시한다. 점 군은 컴퓨터 비전 및 관련 분야에서 기하학적 구조를 표현하기 위한 널리 사용되는 데이...
본 학위논문은 점 군 데이터를 서로 다른 문제 설정에서 활용하는 두 가지 연구 방향을 제시한다. 점 군은 컴퓨터 비전 및 관련 분야에서 기하학적 구조를 표현하기 위한 널리 사용되는 데이터 표현 방식으로 자리 잡고 있다.
첫 번째 연구는 비대응(unpaired) 점 군 완성(point cloud completion) 문제를 다룬다. 이 설정에서는 불완전 점 군과 완전 점 군 사이에 대응관계가 존재하지 않으며, 두 분포가 서로 다를 수 있다. 본 연구에서는 불균형 최적수송(Unbalanced Optimal Transport, UOT) 맵을 기반으로 불완전 점 군 분포와 완전 점 군 분포 사이의 수송 사상을 학습하는 프레임워크 UOT-UPC를 제안한다. 이 접근법은 불완전 점 군과 완전 점 군 사이의 페어(pair) 형태의 지도(supervision) 없이도 점 군 완성을 가능하게 하며, 클래스 불균형 상황에서도 높은 강건성을 바탕으로 경쟁력 있는 성능을 달성한다. 두 번째 연구는 사람 자세(human pose)를 포함하는 텍스트-투-이미지(text-to-image, T2I) 생성 문제에 초점을 둔다. 본 연구에서는 대형 언어 모델(Large Language Model, LLM)을 활용하여 텍스트 프롬프트로부터 점 군 형태의 자세 키포인트를 직접 추론하고, 이를 자세 인지 T2I 모델을 위한 구조적 유도 신호(structural guidance)로 활용하는 방법론 PointT2I를 제안한다. 이 방법론은 LLM 기반 피드백 시스템을 도입하고 자세 특화 학습 없이도, 다양한 도전적 텍스트 프롬프트에 대해 자세가 정확한 이미지를 생성한다.
다국어 초록 (Multilingual Abstract)
This dissertation presents two independent research directions that leverage point cloud data in fundamentally different problem settings. In computer vision and related fields, point clouds are a widely used representation for describing geometric st...
This dissertation presents two independent research directions that leverage point cloud data in fundamentally different problem settings. In computer vision and related fields, point clouds are a widely used representation for describing geometric structure.
The first study addresses unpaired point cloud completion, where incomplete and complete point clouds lack correspondences and may differ in distribution. We introduce UOT-UPC, a framework based on the Unbalanced Optimal Transport (UOT) map that learns a transport mapping between the incomplete and complete point cloud distributions. This approach enables completion without paired supervision and achieves competitive performance with strong robustness to class imbalance. The second study focuses on text-to-image (T2I) generation involving human poses. We propose PointT2I, a framework that uses a Large Language Model to infer pose keypoints—represented as a point cloud—from textual prompts and applies them as structural guidance for pose-aware T2I models. With an LLM-based feedback mechanism and no pose-specific training, the framework generates pose-accurate images across diverse and challenging scenarios.
목차 (Table of Contents)