RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Post-training of Dense Visual Representations : Patch-Level Kernel Alignment for Dense Self-Supervised Learning = 밀집 시각 표현의 사후 학습

    한글로보기

    https://www.riss.kr/link?id=T17450787

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Dense self-supervised learning (SSL) methods showed its effectiveness in enhancing the fine-grained semantic understandings of vision models. However, existing approaches often rely on parametric assumptions or complex post-processing (e.g., clustering, sorting), limiting their flexibility and stability. To overcome these limitations, we introduce Patch-level Kernel Alignment (PaKA), a non-parametric, kernel-based approach that improves the dense representations of pretrained vision encoders with a post-(pre)training. Our method propose a robust and effective alignment objective that captures statistical dependencies which matches the intrinsic structure of high-dimensional dense feature distributions. In addition, we revisit the augmentation strategies inherited from image-level SSL and propose a refined augmentation strategy for dense SSL. Our framework improves dense representations by conducting a lightweight post-training stage on top of a pretrained model. With only 14 hours of additional training on a single GPU, our method achieves state-of-the-art performance across a range of dense vision benchmarks, demonstrating both efficiency and effectiveness.
    번역하기

    Dense self-supervised learning (SSL) methods showed its effectiveness in enhancing the fine-grained semantic understandings of vision models. However, existing approaches often rely on parametric assumptions or complex post-processing (e.g., clusterin...

    Dense self-supervised learning (SSL) methods showed its effectiveness in enhancing the fine-grained semantic understandings of vision models. However, existing approaches often rely on parametric assumptions or complex post-processing (e.g., clustering, sorting), limiting their flexibility and stability. To overcome these limitations, we introduce Patch-level Kernel Alignment (PaKA), a non-parametric, kernel-based approach that improves the dense representations of pretrained vision encoders with a post-(pre)training. Our method propose a robust and effective alignment objective that captures statistical dependencies which matches the intrinsic structure of high-dimensional dense feature distributions. In addition, we revisit the augmentation strategies inherited from image-level SSL and propose a refined augmentation strategy for dense SSL. Our framework improves dense representations by conducting a lightweight post-training stage on top of a pretrained model. With only 14 hours of additional training on a single GPU, our method achieves state-of-the-art performance across a range of dense vision benchmarks, demonstrating both efficiency and effectiveness.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    최근의 밀집자기지도학습 기법은 비전 모델의 정교한 의미 이해 능력을 향상시키는 데 높은 효과를 보여왔다. 그러나 기존 방법들은 특정한 형태를 가정하는 매개변수적 방식이나 군집화·정렬과 같은 복잡한 후처리 절차에 의존하는 경우가 많아, 유연성과 안정성이 제한되는 문제가 존재한다. 본 연구에서는 이러한 한계를 해소하기 위해 패치 단위 커널 정렬 기법(PaKA)을 제안한다. 해당 방법은 비매개변수적이며, 사전학습된 이미지 인코더가 산출하는 밀집 특징 표현을 추가 학습 단계만으로 향상시키는 방법이다. 제안하는 정렬 목적함수는 고차원 특징 분포가 지니는 내재적 구조와 통계적 의존 관계를 효과적으로 포착하여, 더욱 견고하고 효율적인 표현 학습을 가능하게 한다. 또한 기존 이미지 단위 수준 자기지도학습에서 사용되던 데이터 증강 전략을 재검토하고, 밀집 표현 학습에 적합하도록 개선된 증강 절차를 제안한다. 우리의 방법은 사전학습된 모델 위에서 가벼운 추가 학습만 수행해도 밀집 표현 성능을 크게 향상시킨다. 단일 그래픽 연산 장치(GPU)에서 약 14시간의 추가 학습만으로 다양한 밀집 시각 과제에서 최상위 수준의 성능을 달성하여, 효율성과 실효성을 모두 입증하였다.
    번역하기

    최근의 밀집자기지도학습 기법은 비전 모델의 정교한 의미 이해 능력을 향상시키는 데 높은 효과를 보여왔다. 그러나 기존 방법들은 특정한 형태를 가정하는 매개변수적 방식이나 군집화·...

    최근의 밀집자기지도학습 기법은 비전 모델의 정교한 의미 이해 능력을 향상시키는 데 높은 효과를 보여왔다. 그러나 기존 방법들은 특정한 형태를 가정하는 매개변수적 방식이나 군집화·정렬과 같은 복잡한 후처리 절차에 의존하는 경우가 많아, 유연성과 안정성이 제한되는 문제가 존재한다. 본 연구에서는 이러한 한계를 해소하기 위해 패치 단위 커널 정렬 기법(PaKA)을 제안한다. 해당 방법은 비매개변수적이며, 사전학습된 이미지 인코더가 산출하는 밀집 특징 표현을 추가 학습 단계만으로 향상시키는 방법이다. 제안하는 정렬 목적함수는 고차원 특징 분포가 지니는 내재적 구조와 통계적 의존 관계를 효과적으로 포착하여, 더욱 견고하고 효율적인 표현 학습을 가능하게 한다. 또한 기존 이미지 단위 수준 자기지도학습에서 사용되던 데이터 증강 전략을 재검토하고, 밀집 표현 학습에 적합하도록 개선된 증강 절차를 제안한다. 우리의 방법은 사전학습된 모델 위에서 가벼운 추가 학습만 수행해도 밀집 표현 성능을 크게 향상시킨다. 단일 그래픽 연산 장치(GPU)에서 약 14시간의 추가 학습만으로 다양한 밀집 시각 과제에서 최상위 수준의 성능을 달성하여, 효율성과 실효성을 모두 입증하였다.

    더보기

    목차 (Table of Contents)

    • 1. Introduction 1
    • 2. Related Works 5
    • 2.1 Image-level Self-supervised Learning 5
    • 2.2 Dense Self-Supervised Learning 6
    • 2.3 Kernel Alignment 7
    • 1. Introduction 1
    • 2. Related Works 5
    • 2.1 Image-level Self-supervised Learning 5
    • 2.2 Dense Self-Supervised Learning 6
    • 2.3 Kernel Alignment 7
    • 3. Rethinking Distribution Alignment for Dense SSL 8
    • 3.1 Distribution Alignment in Dense Self-Supervised Learning 8
    • 3.2 From Parametric to Non-Parametric Relational Learning 9
    • 4. Our Method 11
    • 4.1 Post-(pre)training Vision Encoders for Dense Representation 11
    • 4.2 Kernel-based Relational Alignment 13
    • 4.3 Formulation of the PaKA Loss 15
    • 4.4 Towards Effective Augmentation Strategies for Dense SSL 16
    • 5. Experiments 20
    • 5.1 Implementation Details 20
    • 5.2 Datasets 21
    • 5.3 Comparison with Prior Works 21
    • 6. Conclusion 29
    • 7. Abstract (In Korean) 40
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼