RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Transformer-based Video Object Segmentation Methods with Target-Specific Object Queries for Eliminating the Association Process = 타겟 특정 객체 쿼리를 사용하여 연관 과정을 제거하는 트랜스포머 기반 비디오 객체 분할 방법

    한글로보기

    https://www.riss.kr/link?id=T17109923

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This dissertation presents novel Transformer-based methods for Video Object Segmentation (VOS) with enhanced Object Queries. VOS is crucial for understanding spatial-temporal dynamics in videos and is valuable in numerous industrial applications, including Semi-supervised Video Object Segmentation (SSVOS) and Video Instance Segmentation (VIS). In recent developments, the Transformer architecture has been effectively adapted for vision-based tasks. The DEtection TRansformer (DETR) introduced Object Queries, which are learnable embeddings representing potential objects within an image. DETR induces the properties of non-maximum suppression (NMS) and region proposal through bipartite matching functions. The flexibility of Object Queries in DETR allows them to induce desired properties by designing specific objective functions. This dissertation leverages this flexibility to enhance performance across different VOS sub-tasks, removing the need for auxiliary association processes. First, I propose an object-centric VOS (OCVOS) method for SSVOS using query-based Transformer decoder blocks to extract object-wise information. This method introduces target-specific queries that segment specified objects directly, eliminating the need for auxiliary association processes for tracking. Building on this, I extend the task to VIS and introduce VideoMaskDINO, an extension of MaskDINO. This method uses the Denoising (DN) task from DN-DETR for video tracking, sharing principles with OCVOS. To enhance tracking reliability, I modify the training schemes to simulate tracking scenarios more accurately and initialize the DN parts’ anchors with Kalman filter predictions. This approach demonstrates consistent tracking ability with frame-level learning alone. Finally, I propose a novel offline VIS approach by integrating upsampling layers into the Transformer decoder. Existing methods face challenges with query capacity and ID-switching. By introducing upsampling of video queries along the temporal axis, I facilitate association properties and refine predictions. This method handles video queries with arbitrary temporal dimensions and shows notable performance, particularly in long and challenging videos. Overall, this dissertation demonstrates the adaptability and robustness of the Transformer architecture with enhanced Object Queries across various VOS tasks, providing comprehensive solutions for improving video segmentation performance and eliminating the auxiliary association processes.
    번역하기

    This dissertation presents novel Transformer-based methods for Video Object Segmentation (VOS) with enhanced Object Queries. VOS is crucial for understanding spatial-temporal dynamics in videos and is valuable in numerous industrial applications, incl...

    This dissertation presents novel Transformer-based methods for Video Object Segmentation (VOS) with enhanced Object Queries. VOS is crucial for understanding spatial-temporal dynamics in videos and is valuable in numerous industrial applications, including Semi-supervised Video Object Segmentation (SSVOS) and Video Instance Segmentation (VIS). In recent developments, the Transformer architecture has been effectively adapted for vision-based tasks. The DEtection TRansformer (DETR) introduced Object Queries, which are learnable embeddings representing potential objects within an image. DETR induces the properties of non-maximum suppression (NMS) and region proposal through bipartite matching functions. The flexibility of Object Queries in DETR allows them to induce desired properties by designing specific objective functions. This dissertation leverages this flexibility to enhance performance across different VOS sub-tasks, removing the need for auxiliary association processes. First, I propose an object-centric VOS (OCVOS) method for SSVOS using query-based Transformer decoder blocks to extract object-wise information. This method introduces target-specific queries that segment specified objects directly, eliminating the need for auxiliary association processes for tracking. Building on this, I extend the task to VIS and introduce VideoMaskDINO, an extension of MaskDINO. This method uses the Denoising (DN) task from DN-DETR for video tracking, sharing principles with OCVOS. To enhance tracking reliability, I modify the training schemes to simulate tracking scenarios more accurately and initialize the DN parts’ anchors with Kalman filter predictions. This approach demonstrates consistent tracking ability with frame-level learning alone. Finally, I propose a novel offline VIS approach by integrating upsampling layers into the Transformer decoder. Existing methods face challenges with query capacity and ID-switching. By introducing upsampling of video queries along the temporal axis, I facilitate association properties and refine predictions. This method handles video queries with arbitrary temporal dimensions and shows notable performance, particularly in long and challenging videos. Overall, this dissertation demonstrates the adaptability and robustness of the Transformer architecture with enhanced Object Queries across various VOS tasks, providing comprehensive solutions for improving video segmentation performance and eliminating the auxiliary association processes.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    이 논문은 향상된 오브젝트 쿼리를 활용한 Transformer 기반의 새로운 비디오 객체 분할(Video Object Segmentation, VOS) 방법을 제시한다. VOS는 비디오의 시공간 동역학을 이해하는 데 중요하며, 반자동 비디오 객체 분할(Semi-supervised Video
    Object Segmentation, SSVOS) 및 비디오 인스턴스 분할(Video Instance Segmentation, VIS) 등 다양한 산업 응용 분야에서 가치를 지닌다. 최근 자연어 처리 분야에서 큰 성과를 거둔 Transformer 구조는 비전 기반 작업에 효과적으로 적용되었다. 특히 DEtection TRansformer(DETR)는 이미지 내의 잠재적 객체를 나타내는 학습 가능한 임베딩인 오브젝트 쿼리(object query)를 도입했다. DETR은 오브젝트 쿼리에 이분 매칭(bipartite matching)을 통해 비-최대 억제(Non-Maximum Suppression, NMS)와 영역 제안(region proposal)의 특성을 유도되는 것을 확인하였으며, 특정 목표 함수를 설계하여 원하는 특성을 유도할 수 있는 유연성을 지니고 있다는 걸 보였다. 본 논문에서는 이 유연성을 활용하여 다양한 VOS 하위 작업에서 성능을 향상시키고 추가적인 연관(Association) 과정의 필요성을 제거하는 방법을 제안한다. 먼저, 본 논문은 SSVOS를 위한 오브젝트 중심 VOS(OCVOS) 방법을 제안한다. 이 방법은 쿼리 기반 Transformer 디코더 블록을 사용하여 객체 단위의 정보를 추출한다. 또한 이 방법은 특정 객체를 직접 분할하는 타겟-특정 쿼리를 도입하여 추적을 위한 연관(association) 과정의 필요성을 제거한다. 이를 바탕으로, 온라인 VIS 작업으로 확장하고 MaskDINO의 확장판인 VideoMaskDINO를 제안한다. 이 방법은 OCVOS와 원리를 공유하는 DN-DETR의 디노이징(Denoising, DN) 작업을 비디오 추적에 사용한다. 추적 신뢰성을 향상시키기 위해, 추적 시나리오를 보다 정확하게 시뮬레이션하도록 학습 방식을 수정하고, 칼만 필터 예측으로 DN 부분의 앵커를 초기화한다. 이 접근 방식은 프레임 단위 학습만으로도 일관된 추적 능력을 보여준다. 마지막으로, Transformer 디코더에 업샘플링 레이어를 통합하여 오프라인 VIS를 위한 새로운 접근 방식을 제안한다. 기존 방법은 작은 쿼리 용량으로 인한 ID 스위칭 문제에 직면해 있다. 비디오 쿼리를 시간 축을 따라 업샘플링하여 연관 속성을 촉진하고 더욱 정제된 결과를 내도록 한다. 이 방법은 임의의 시간적 차원을 가진
    비디오 쿼리를 처리할 수 있으며, 특히 길고 도전적인 비디오에서 눈에 띄는 성능을 보여준다. 결론적으로, 이 논문은 다양한 VOS 작업에서 향상된 오브젝트 쿼리를 사용한 Transformer 아키텍처의 적응성과 견고성을 입증하며, 추가적인 연관 과정을 제거
    하여 비디오 분할 성능을 개선하기 위한 포괄적인 솔루션을 제공한다.
    번역하기

    이 논문은 향상된 오브젝트 쿼리를 활용한 Transformer 기반의 새로운 비디오 객체 분할(Video Object Segmentation, VOS) 방법을 제시한다. VOS는 비디오의 시공간 동역학을 이해하는 데 중요하며, 반자...

    이 논문은 향상된 오브젝트 쿼리를 활용한 Transformer 기반의 새로운 비디오 객체 분할(Video Object Segmentation, VOS) 방법을 제시한다. VOS는 비디오의 시공간 동역학을 이해하는 데 중요하며, 반자동 비디오 객체 분할(Semi-supervised Video
    Object Segmentation, SSVOS) 및 비디오 인스턴스 분할(Video Instance Segmentation, VIS) 등 다양한 산업 응용 분야에서 가치를 지닌다. 최근 자연어 처리 분야에서 큰 성과를 거둔 Transformer 구조는 비전 기반 작업에 효과적으로 적용되었다. 특히 DEtection TRansformer(DETR)는 이미지 내의 잠재적 객체를 나타내는 학습 가능한 임베딩인 오브젝트 쿼리(object query)를 도입했다. DETR은 오브젝트 쿼리에 이분 매칭(bipartite matching)을 통해 비-최대 억제(Non-Maximum Suppression, NMS)와 영역 제안(region proposal)의 특성을 유도되는 것을 확인하였으며, 특정 목표 함수를 설계하여 원하는 특성을 유도할 수 있는 유연성을 지니고 있다는 걸 보였다. 본 논문에서는 이 유연성을 활용하여 다양한 VOS 하위 작업에서 성능을 향상시키고 추가적인 연관(Association) 과정의 필요성을 제거하는 방법을 제안한다. 먼저, 본 논문은 SSVOS를 위한 오브젝트 중심 VOS(OCVOS) 방법을 제안한다. 이 방법은 쿼리 기반 Transformer 디코더 블록을 사용하여 객체 단위의 정보를 추출한다. 또한 이 방법은 특정 객체를 직접 분할하는 타겟-특정 쿼리를 도입하여 추적을 위한 연관(association) 과정의 필요성을 제거한다. 이를 바탕으로, 온라인 VIS 작업으로 확장하고 MaskDINO의 확장판인 VideoMaskDINO를 제안한다. 이 방법은 OCVOS와 원리를 공유하는 DN-DETR의 디노이징(Denoising, DN) 작업을 비디오 추적에 사용한다. 추적 신뢰성을 향상시키기 위해, 추적 시나리오를 보다 정확하게 시뮬레이션하도록 학습 방식을 수정하고, 칼만 필터 예측으로 DN 부분의 앵커를 초기화한다. 이 접근 방식은 프레임 단위 학습만으로도 일관된 추적 능력을 보여준다. 마지막으로, Transformer 디코더에 업샘플링 레이어를 통합하여 오프라인 VIS를 위한 새로운 접근 방식을 제안한다. 기존 방법은 작은 쿼리 용량으로 인한 ID 스위칭 문제에 직면해 있다. 비디오 쿼리를 시간 축을 따라 업샘플링하여 연관 속성을 촉진하고 더욱 정제된 결과를 내도록 한다. 이 방법은 임의의 시간적 차원을 가진
    비디오 쿼리를 처리할 수 있으며, 특히 길고 도전적인 비디오에서 눈에 띄는 성능을 보여준다. 결론적으로, 이 논문은 다양한 VOS 작업에서 향상된 오브젝트 쿼리를 사용한 Transformer 아키텍처의 적응성과 견고성을 입증하며, 추가적인 연관 과정을 제거
    하여 비디오 분할 성능을 개선하기 위한 포괄적인 솔루션을 제공한다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents iii
    • List of Tables vi
    • List of Figures viii
    • 1 Introduction 1
    • Abstract i
    • Contents iii
    • List of Tables vi
    • List of Figures viii
    • 1 Introduction 1
    • 1.1 Motivation 1
    • 1.2 Contribution . 2
    • 1.3 Contents 4
    • 2 Related Works 5
    • 2.1 DEtection TRansformer (DETR) 5
    • 2.1.1 Model Architecture 5
    • 2.1.2 Training Process 7
    • 2.1.3 Discussion of Object Queries 8
    • 3 OCVOS: Object-Centric Representation for Video Object Segmentation 11
    • 3.1 Motivation and Overview 11
    • 3.2 Method 14
    • 3.2.1 Memory Readout Module . 14
    • iii
    • 3.2.2 Transformer Encoder . 14
    • 3.2.3 Transformer Decoder . 15
    • 3.3 Experimental Results 17
    • 3.3.1 Implementation Details 17
    • 3.3.2 Performance on Benchmarks 17
    • 3.3.3 Ablation Studies 19
    • 3.4 Summary 21
    • 4 VideoMaskDINO: MaskDINO for Video Instance Segmentation 22
    • 4.1 Motivation and Overview 22
    • 4.2 Related Works 23
    • 4.2.1 DN-DETR 23
    • 4.2.2 Kalman Filter 26
    • 4.3 Method 28
    • 4.3.1 MaskDINO better than Mask2Former? 28
    • 4.3.2 DN part tracking ability 30
    • 4.3.3 Issue on DN part 34
    • 4.4 Expermental Results 39
    • 4.4.1 Datasets and Metrics 39
    • 4.4.2 Ablation Study . 40
    • 4.5 Summary 43
    • 5 UPVIS: Upsampled Video Query for Offline Video Instance Segmentation 48
    • 5.1 Motivation and Overview 48
    • 5.2 Related Works 52
    • 5.2.1 Online Video Instance Segmentation 52
    • 5.2.2 Offline Video Instance Segmentation . 53
    • 5.3 Method 54
    • 5.3.1 Frame-level Detector . 54
    • iv
    • 5.3.2 Object Token Encoder 56
    • 5.3.3 Object Token Decoder 56
    • 5.3.4 Prediction . 59
    • 5.3.5 Training 60
    • 5.4 Experimental Results 60
    • 5.4.1 Datasets and Metrics 60
    • 5.4.2 Implementation Details 61
    • 5.4.3 Performance on Benchmarks 64
    • 5.4.4 Behavior of video query 68
    • 5.4.5 Ablation Study . 69
    • 5.4.6 Auxiliary Ablation Study 72
    • 5.5 Limitation 73
    • 5.6 Summary 74
    • 6 Conclusion 81
    • Bibliography 83
    • Abstract (In Korean) 91
    • v
    더보기

    참고문헌 (Reference)

    1. Fast r-cnn, R. Girshick, in Proceedings of the IEEE international conference on computer vision pp. 1440–1448, , 2015

    2. Mask r-cnn, K. He, G. Gkioxari, R. Girshick, P. Doll´ar and, in Proceedings of the IEEE international conference on computer vision pp. 2961–2969, , 2017

    3. Attention is all you need, A. N. Gomez, I. Polosukhin, A. Vaswani, N. Shazeer, L. Jones, J. Uszkoreit, Ł. Kaiser and, N. Parmar, Advances in neural information processing systems, vol. 30, , 2017

    4. Attention is all you need, in, A. N. Gomez, Ł. Kaiser and, I. Polosukhin, N. Parmar, L. Jones, A. Vaswani, J. Uszkoreit, N. Shazeer, Advances in Neural Information Processing Systems, pp. 5998–6008, , 2017

    5. Tracking instances as queries,, S. Yang, Y. Shan, W. Liu, B. Feng and, Y. Li, Y. Fang, X. Wang, arXiv preprint arXiv:2106.11963, , 2021

    6. Video instance segmentation, in, Y. Fan and, L. Yang, N. Xu, Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 5188–5197, , 2019

    7. Microsoft coco: Common objects in context, D. Ramanan, P. Doll´ar and, T.-Y. Lin, P. Perona, J. Hays, S. Belongie, C. L. Zitnick, M. Maire, Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, Proceedings, Part V 13. Springer, 2014, pp. 740–755., , 2014

    8. Relation networks for object detection, in, Y. Wei, J. Dai and, J. Gu, H. Hu, Z. Zhang, Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3588–3597, , 2018

    9. Mask2former for video instance segmentation, B. Cheng, R. Girdhar and, A. Choudhuri, A. Kirillov, A. G. Schwing, I. Misra, arXiv preprint arXiv:2112.10764, , 2021

    10. Vision transformer with deformable attention,, G. Huang, Z. Xia, X. Pan, L. E. Li and, S. Song, in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4794–4803, , 2022

    1. Fast r-cnn, R. Girshick, in Proceedings of the IEEE international conference on computer vision pp. 1440–1448, , 2015

    2. Mask r-cnn, K. He, G. Gkioxari, R. Girshick, P. Doll´ar and, in Proceedings of the IEEE international conference on computer vision pp. 2961–2969, , 2017

    3. Attention is all you need, A. N. Gomez, I. Polosukhin, A. Vaswani, N. Shazeer, L. Jones, J. Uszkoreit, Ł. Kaiser and, N. Parmar, Advances in neural information processing systems, vol. 30, , 2017

    4. Attention is all you need, in, A. N. Gomez, Ł. Kaiser and, I. Polosukhin, N. Parmar, L. Jones, A. Vaswani, J. Uszkoreit, N. Shazeer, Advances in Neural Information Processing Systems, pp. 5998–6008, , 2017

    5. Tracking instances as queries,, S. Yang, Y. Shan, W. Liu, B. Feng and, Y. Li, Y. Fang, X. Wang, arXiv preprint arXiv:2106.11963, , 2021

    6. Video instance segmentation, in, Y. Fan and, L. Yang, N. Xu, Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 5188–5197, , 2019

    7. Microsoft coco: Common objects in context, D. Ramanan, P. Doll´ar and, T.-Y. Lin, P. Perona, J. Hays, S. Belongie, C. L. Zitnick, M. Maire, Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, Proceedings, Part V 13. Springer, 2014, pp. 740–755., , 2014

    8. Relation networks for object detection, in, Y. Wei, J. Dai and, J. Gu, H. Hu, Z. Zhang, Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3588–3597, , 2018

    9. Mask2former for video instance segmentation, B. Cheng, R. Girdhar and, A. Choudhuri, A. Kirillov, A. G. Schwing, I. Misra, arXiv preprint arXiv:2112.10764, , 2021

    10. Vision transformer with deformable attention,, G. Huang, Z. Xia, X. Pan, L. E. Li and, S. Song, in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4794–4803, , 2022

    11. Deep residual learning for image recognition in, K. He, S. Ren and, X. Zhang, J. Sun, Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, , 2016

    12. End-to-end object detection with transformers, in, A. Kirillov and, N. Usunier, N. Carion, S. Zagoruyko, G. Synnaeve, F. Massa, ECCV, , 2020

    13. Feature pyramid networks for object detection, in, S. Belongie, R. Girshick, T.-Y. Lin, P. Doll´ar, K. He, B. Hariharan and, Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2117–2125, , 2017

    14. Occluded video instance segmentation: A benchmark,, S. Belongie, P. H. Torr and, Y. Hu, X. Liu, S. Bai, J. Qi, A. Yuille, Y. Gao, X. Wang, X. Bai, International Journal of Computer Vision, vol. 130, no. 8, pp. 2022–2039,, , 2022

    15. The 2017 davis challenge on video object segmentation, L. Van Gool, A. Sorkine-Hornung and, J. Pont-Tuset, P. Arbel´aez, S. Caelles, F. Perazzi, arXiv preprint arXiv:1704.00675, , 2017

    16. Transvos: Video object segmentation with transformers, J. Mei, M. Wang, Y. Liu, Y. Yuan and, Y. Lin, arXiv e-prints, pp. arXiv–2106, , 2021

    17. A generalized framework for video instance segmentation, M. Heo, J.-Y. Lee and, H. Kim, S. J. Kim, J. Hyun, S. W. Oh, S. Hwang, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 14 623–14 632, , 2023

    18. Videomatch: Matching based video object segmentation in, J.-B. Huang and, Y.-T. Hu, A. G. Schwing, Proceedings of the European conference on computer vision (ECCV), pp. 54–70, , 2018

    19. Dvis: Decoupled video instance segmentation framework in, X. Tian, S. Ji, Y. Zhang and, T. Zhang, P. Wan, Y. Wu, X. Wang, Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 1282– 1291, , 2023

    20. End-to-end video instance segmentation with transformers, Z. Xu, X. Wang, H. Xia, B. Cheng, C. Shen, Y. Wang, H. Shen and, in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), , 2021

    21. U-netConvolutional networks for biomedical image segmentation, O. Ronneberger, P. Fischer and, T. Brox, in Medical Image Computing and Computer- Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, Proceedings, Part III 18. Springer, pp. 234–241, , 2015

    22. Video object segmentation using space-time memory networks, in, J.-Y. Lee, N. Xu and, S. W. Oh, S. J. Kim, Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9226–9235, , 2019

    23. In defense of online models for video instance segmentation, in, S. Bai, Q. Liu, J. Wu, A. Yuille and, X. Bai, Y. Jiang, European Conference on Computer Vision. Springer, pp. 588–605, , 2022

    24. YouTube-VOS: A large-scale video object segmentation benchmark,, T. Huang, Y. Liang, Y. Fan, L. Yang, J. Yang and, D. Yue, N. Xu, arXiv preprint arXiv:1809.03327, , 2018

    25. Tcovis: Temporally consistent online video instance segmentation, B. Yu, J. Lu, Y. Rao, J. Li, J. Zhou and, in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 1097–1107, , 2023

    26. Universal instance perception as object discovery and retrieval,, P. Luo, J. Wu, H. Lu, Z. Yuan and, B. Yan, Y. Jiang, D. Wang, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 15 325– 15 336, , 2023

    27. Ctvis: Consistent training for online video instance segmentation, C. Fan, Y. Zhuge and, W. Mao, Q. Zhong, C. Shen, Z. Wang, Y. Liu, L. Y. Wu, H. Chen, K. Ying, in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 899–908, , 2023

    28. Enhanced deep residual networks for single image super-resolution, S. Nah and, K. Mu Lee, B. Lim, H. Kim, S. Son, in Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pp. 136–144, , 2017

    29. Crossover learning for fast online video instance segmentation, in, S. Yang, Y. Li, C. Fang, Y. Shan, X. Wang, Y. Fang, B. Feng and, W. Liu, Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 8043–8052, , 2021

    30. Maskedattention mask transformer for universal image segmentation,, A. Kirillov and, B. Cheng, I. Misra, R. Girdhar, A. G. Schwing, in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), , 2022

    31. Vita: Video instance segmentation via object token association, in, M. Heo, S. W. Oh, J.-Y. Lee and, S. J. Kim, S. Hwang, Advances in Neural Information Processing Systems, , 2022

    32. Associating objects with transformers for video object segmentation, Y. Yang, Z. Yang, Y. Wei and, in Advances in Neural Information Processing Systems (NeurIPS), , 2021

    33. Dn-detr: Accelerate detr training by introducing query denoising, in, S. Liu, J. Guo, F. Li, H. Zhang, L. Zhang, L. M. Ni and, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13 619–13 627, , 2022

    34. Seqformer: Sequential transformer for video instance segmentation in, W. Zhang and, Y. Jiang, S. Bai, X. Bai, J. Wu, ECCV, , 2022

    35. Dvis++: Improved decoupled framework for universal video segmentation, S. Ji, X. Tian, P. Wan, X. Wang, X. Tao, Y. Zhou, Z. Wang and, T. Zhang, Y. Zhang, Y. Wu, arXiv preprint arXiv:2312.13305, 2023, , 2023

    36. Novis: A case for end-to-end near-online video instance segmentation,, Y. Fan, M. Feiszli, T. Meinhardt, R. Ranjan, L. Leal-Taixe and, arXiv preprint arXiv:2308.15266, 2023, , 2023

    37. Joint inductive and transductive learning for video object segmentation, H. Li, Y. Mao, N. Wang, W. Zhou and, in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), , 2021

    38. Deformable detr: Deformable transformers for end-to-end object detection, J. Dai, X. Wang and, X. Zhu, B. Li, L. Lu, W. Su, arXiv preprint arXiv:2010.04159, , 2020

    39. Sstvos: Sparse spatiotemporal transformers for video object segmentation, A. Ahmed, G. W. Taylor, B. Duke, P. Aarabi and, C. Wolf, in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), , 2021

    40. FEELVOS: Fast end-to-end embedding learning for video object segmentation, F. Schroff, Y. Chai, L.-C. Chen, H. Adam, B. Leibe and, P. Voigtlaender, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9473–9482, , 2019

    41. Refinevis: Video instance segmentation with temporal attention refinement, Q. You and, P. Chu, A. Abrantes, Z. Liu, J. Wang, arXiv preprint arXiv:2306.04774, 2023, , 2023

    42. Video instance segmentation using inter-frame communication transformers,, S. Hwang, S. W. Oh and, M. Heo, S. J. Kim, Advances in Neural Information Processing Systems, vol. 34, pp. 13 352–13 363, , 2021

    43. An image is worth 16x16 words: Transformers for image recognition at scale, A. Kolesnikov, S. Gelly, M. Dehghani, A. Dosovitskiy, L. Beyer, D. Weissenborn, X. Zhai, T. Unterthiner, N. Houlsby, M. Minderer, G. Heigold, J. Uszkoreit and, International Conference on Learning Representations, 2021. [Online]. Available: https://openreview. net/forum?id=YicbFdNTTy, , 2021

    44. Swin transformer: Hierarchical vision transformer using shifted windows, in, Y. Cao, Z. Zhang, Z. Liu, Y. Wei, H. Hu, S. Lin and, B. Guo, Y. Lin, Proceedings of the IEEE/CVF international conference on computer vision, pp. 10 012– 10 022, , 2021

    45. Temporally efficient vision transformer for video instance segmentation, in, J. Fang, Y. Fang, X. Zhao and, X. Wang, Y. Shan, Y. Li, W. Liu, S. Yang, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2885–2895, , 2022

    46. Collaborative video object segmentation by foreground-background integration, in, Y. Yang, Y. Wei and, Z. Yang, European Conference on Computer Vision. Springer, pp. 332–348, , 2020

    47. Video object segmentation using kernelized memory network with multiple kernels,, E. Kim, H. Seong, J. Hyun and, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, pp. 2595–2612, , 2022

    48. Learning position and target consistency for memory-based video object segmentation, P. Pan, Y. Xu and, R. Jin, L. Hu, B. Zhang, P. Zhang, in IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, Computer Vision Foundation / IEEE, 2021, pp. 4144–4154Online Available: https://openaccess. thecvf. com/content/CVPR2021/html/ Hu\ Learning\ Position\ and\ Target\ Consistency\ for\ Memory-Based\ Video\ Object\ Segmentation\ CVPR\ 2021\ paper. html, , 2021

    49. VNet: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation, N. Navab and, S.A. Ahmadi, F. Milletari, 2016 fourth international conference on 3D vision (3DV). Ieee, pp. 565–571, , 2016

    50. Gratt-vis: Gated residual attention for auto rectifying video instance segmentation,, B. Menze, T. Hannan, R. Koner, M. Schubert and, S. Shit, T. Seidl, V. Tresp, M. Bernhard, arXiv preprint arXiv:2305.17096, 2023, , 2023

    51. Minvis: A minimal video instance segmentation framework without video-based training, A. Anandkumar, D.-A. Huang, Z. Yu and, Advances in Neural Information Processing Systems, vol. 35, pp. 31 265–31 277, , 2022

    52. Classifying, segmenting, and tracking object instances in video with mask propagation, in, L. Torresani, G. Bertasius and, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9739–9748, , 2020

    53. Sipmask: Spatial information preservation for fast image and video instance segmentation, in, F. S. Khan, J. Cao, Y. Pang and, R. M. Anwer, H. Cholakkal, L. Shao, Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK,, Proceedings, Part XIV 16. Springer, 2020, pp. 1–18, , 2020

    54. Visolo: Grid-based space-time aggregation for efficient online video instance segmentation in, S. J. Kim, Y. Park, H. Kim, S. H. Han, M.-J. Kim and, S. Hwang, S. W. Oh, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2896–2905, , 2022

    55. Mask dino: Towards a unified transformer-based framework for object detection and segmentation,, F. Li, L. Zhang, H. Xu, H.-Y. Shum, H. Zhang, S. Liu, L. M. Ni and, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 3041–3050, , 2023

    56. Rethinking space-time networks with improved memory coverage for efficient video object segmentation,, H. K. Cheng, C.-K. Tang, Y.-W. Tai and, Advances in Neural Information Processing Systems, vol. 34, pp. 11 781–11 794, , 2021

    57. Video object segmentation with adaptive feature bank and uncertain-region refinement in M. Ranzato R., X. Li, Y. Liang, N. Jafari and, J. Chen, M., Advances in Neural Information Processing Systems H. Larochelle Hadsell F. Balcan and H. Lin Eds. vol. 33. Curran Associates, Inc., pp. 3430–3441. [Online]. Available: https://proceedings. neurips. cc/paper/2020/file/ 234833147b97bb6aed53a8f4f1c7a7d8-Paper. pdf, , 2020

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼