1. Fast r-cnn, R. Girshick, in Proceedings of the IEEE international conference on computer vision pp. 1440–1448, , 2015
2. Mask r-cnn, K. He, G. Gkioxari, R. Girshick, P. Doll´ar and, in Proceedings of the IEEE international conference on computer vision pp. 2961–2969, , 2017
3. Attention is all you need, A. N. Gomez, I. Polosukhin, A. Vaswani, N. Shazeer, L. Jones, J. Uszkoreit, Ł. Kaiser and, N. Parmar, Advances in neural information processing systems, vol. 30, , 2017
4. Attention is all you need, in, A. N. Gomez, Ł. Kaiser and, I. Polosukhin, N. Parmar, L. Jones, A. Vaswani, J. Uszkoreit, N. Shazeer, Advances in Neural Information Processing Systems, pp. 5998–6008, , 2017
5. Tracking instances as queries,, S. Yang, Y. Shan, W. Liu, B. Feng and, Y. Li, Y. Fang, X. Wang, arXiv preprint arXiv:2106.11963, , 2021
6. Video instance segmentation, in, Y. Fan and, L. Yang, N. Xu, Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 5188–5197, , 2019
7. Microsoft coco: Common objects in context, D. Ramanan, P. Doll´ar and, T.-Y. Lin, P. Perona, J. Hays, S. Belongie, C. L. Zitnick, M. Maire, Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, Proceedings, Part V 13. Springer, 2014, pp. 740–755., , 2014
8. Relation networks for object detection, in, Y. Wei, J. Dai and, J. Gu, H. Hu, Z. Zhang, Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3588–3597, , 2018
9. Mask2former for video instance segmentation, B. Cheng, R. Girdhar and, A. Choudhuri, A. Kirillov, A. G. Schwing, I. Misra, arXiv preprint arXiv:2112.10764, , 2021
10. Vision transformer with deformable attention,, G. Huang, Z. Xia, X. Pan, L. E. Li and, S. Song, in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4794–4803, , 2022
11. Deep residual learning for image recognition in, K. He, S. Ren and, X. Zhang, J. Sun, Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, , 2016
12. End-to-end object detection with transformers, in, A. Kirillov and, N. Usunier, N. Carion, S. Zagoruyko, G. Synnaeve, F. Massa, ECCV, , 2020
13. Feature pyramid networks for object detection, in, S. Belongie, R. Girshick, T.-Y. Lin, P. Doll´ar, K. He, B. Hariharan and, Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2117–2125, , 2017
14. Occluded video instance segmentation: A benchmark,, S. Belongie, P. H. Torr and, Y. Hu, X. Liu, S. Bai, J. Qi, A. Yuille, Y. Gao, X. Wang, X. Bai, International Journal of Computer Vision, vol. 130, no. 8, pp. 2022–2039,, , 2022
15. The 2017 davis challenge on video object segmentation, L. Van Gool, A. Sorkine-Hornung and, J. Pont-Tuset, P. Arbel´aez, S. Caelles, F. Perazzi, arXiv preprint arXiv:1704.00675, , 2017
16. Transvos: Video object segmentation with transformers, J. Mei, M. Wang, Y. Liu, Y. Yuan and, Y. Lin, arXiv e-prints, pp. arXiv–2106, , 2021
17. A generalized framework for video instance segmentation, M. Heo, J.-Y. Lee and, H. Kim, S. J. Kim, J. Hyun, S. W. Oh, S. Hwang, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 14 623–14 632, , 2023
18. Videomatch: Matching based video object segmentation in, J.-B. Huang and, Y.-T. Hu, A. G. Schwing, Proceedings of the European conference on computer vision (ECCV), pp. 54–70, , 2018
19. Dvis: Decoupled video instance segmentation framework in, X. Tian, S. Ji, Y. Zhang and, T. Zhang, P. Wan, Y. Wu, X. Wang, Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 1282– 1291, , 2023
20. End-to-end video instance segmentation with transformers, Z. Xu, X. Wang, H. Xia, B. Cheng, C. Shen, Y. Wang, H. Shen and, in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), , 2021
21. U-netConvolutional networks for biomedical image segmentation, O. Ronneberger, P. Fischer and, T. Brox, in Medical Image Computing and Computer- Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, Proceedings, Part III 18. Springer, pp. 234–241, , 2015
22. Video object segmentation using space-time memory networks, in, J.-Y. Lee, N. Xu and, S. W. Oh, S. J. Kim, Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9226–9235, , 2019
23. In defense of online models for video instance segmentation, in, S. Bai, Q. Liu, J. Wu, A. Yuille and, X. Bai, Y. Jiang, European Conference on Computer Vision. Springer, pp. 588–605, , 2022
24. YouTube-VOS: A large-scale video object segmentation benchmark,, T. Huang, Y. Liang, Y. Fan, L. Yang, J. Yang and, D. Yue, N. Xu, arXiv preprint arXiv:1809.03327, , 2018
25. Tcovis: Temporally consistent online video instance segmentation, B. Yu, J. Lu, Y. Rao, J. Li, J. Zhou and, in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 1097–1107, , 2023
26. Universal instance perception as object discovery and retrieval,, P. Luo, J. Wu, H. Lu, Z. Yuan and, B. Yan, Y. Jiang, D. Wang, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 15 325– 15 336, , 2023
27. Ctvis: Consistent training for online video instance segmentation, C. Fan, Y. Zhuge and, W. Mao, Q. Zhong, C. Shen, Z. Wang, Y. Liu, L. Y. Wu, H. Chen, K. Ying, in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 899–908, , 2023
28. Enhanced deep residual networks for single image super-resolution, S. Nah and, K. Mu Lee, B. Lim, H. Kim, S. Son, in Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pp. 136–144, , 2017
29. Crossover learning for fast online video instance segmentation, in, S. Yang, Y. Li, C. Fang, Y. Shan, X. Wang, Y. Fang, B. Feng and, W. Liu, Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 8043–8052, , 2021
30. Maskedattention mask transformer for universal image segmentation,, A. Kirillov and, B. Cheng, I. Misra, R. Girdhar, A. G. Schwing, in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), , 2022
31. Vita: Video instance segmentation via object token association, in, M. Heo, S. W. Oh, J.-Y. Lee and, S. J. Kim, S. Hwang, Advances in Neural Information Processing Systems, , 2022
32. Associating objects with transformers for video object segmentation, Y. Yang, Z. Yang, Y. Wei and, in Advances in Neural Information Processing Systems (NeurIPS), , 2021
33. Dn-detr: Accelerate detr training by introducing query denoising, in, S. Liu, J. Guo, F. Li, H. Zhang, L. Zhang, L. M. Ni and, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13 619–13 627, , 2022
34. Seqformer: Sequential transformer for video instance segmentation in, W. Zhang and, Y. Jiang, S. Bai, X. Bai, J. Wu, ECCV, , 2022
35. Dvis++: Improved decoupled framework for universal video segmentation, S. Ji, X. Tian, P. Wan, X. Wang, X. Tao, Y. Zhou, Z. Wang and, T. Zhang, Y. Zhang, Y. Wu, arXiv preprint arXiv:2312.13305, 2023, , 2023
36. Novis: A case for end-to-end near-online video instance segmentation,, Y. Fan, M. Feiszli, T. Meinhardt, R. Ranjan, L. Leal-Taixe and, arXiv preprint arXiv:2308.15266, 2023, , 2023
37. Joint inductive and transductive learning for video object segmentation, H. Li, Y. Mao, N. Wang, W. Zhou and, in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), , 2021
38. Deformable detr: Deformable transformers for end-to-end object detection, J. Dai, X. Wang and, X. Zhu, B. Li, L. Lu, W. Su, arXiv preprint arXiv:2010.04159, , 2020
39. Sstvos: Sparse spatiotemporal transformers for video object segmentation, A. Ahmed, G. W. Taylor, B. Duke, P. Aarabi and, C. Wolf, in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), , 2021
40. FEELVOS: Fast end-to-end embedding learning for video object segmentation, F. Schroff, Y. Chai, L.-C. Chen, H. Adam, B. Leibe and, P. Voigtlaender, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9473–9482, , 2019
41. Refinevis: Video instance segmentation with temporal attention refinement, Q. You and, P. Chu, A. Abrantes, Z. Liu, J. Wang, arXiv preprint arXiv:2306.04774, 2023, , 2023
42. Video instance segmentation using inter-frame communication transformers,, S. Hwang, S. W. Oh and, M. Heo, S. J. Kim, Advances in Neural Information Processing Systems, vol. 34, pp. 13 352–13 363, , 2021
43. An image is worth 16x16 words: Transformers for image recognition at scale, A. Kolesnikov, S. Gelly, M. Dehghani, A. Dosovitskiy, L. Beyer, D. Weissenborn, X. Zhai, T. Unterthiner, N. Houlsby, M. Minderer, G. Heigold, J. Uszkoreit and, International Conference on Learning Representations, 2021. [Online]. Available: https://openreview. net/forum?id=YicbFdNTTy, , 2021
44. Swin transformer: Hierarchical vision transformer using shifted windows, in, Y. Cao, Z. Zhang, Z. Liu, Y. Wei, H. Hu, S. Lin and, B. Guo, Y. Lin, Proceedings of the IEEE/CVF international conference on computer vision, pp. 10 012– 10 022, , 2021
45. Temporally efficient vision transformer for video instance segmentation, in, J. Fang, Y. Fang, X. Zhao and, X. Wang, Y. Shan, Y. Li, W. Liu, S. Yang, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2885–2895, , 2022
46. Collaborative video object segmentation by foreground-background integration, in, Y. Yang, Y. Wei and, Z. Yang, European Conference on Computer Vision. Springer, pp. 332–348, , 2020
47. Video object segmentation using kernelized memory network with multiple kernels,, E. Kim, H. Seong, J. Hyun and, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, pp. 2595–2612, , 2022
48. Learning position and target consistency for memory-based video object segmentation, P. Pan, Y. Xu and, R. Jin, L. Hu, B. Zhang, P. Zhang, in IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, Computer Vision Foundation / IEEE, 2021, pp. 4144–4154Online Available: https://openaccess. thecvf. com/content/CVPR2021/html/ Hu\ Learning\ Position\ and\ Target\ Consistency\ for\ Memory-Based\ Video\ Object\ Segmentation\ CVPR\ 2021\ paper. html, , 2021
49. VNet: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation, N. Navab and, S.A. Ahmadi, F. Milletari, 2016 fourth international conference on 3D vision (3DV). Ieee, pp. 565–571, , 2016
50. Gratt-vis: Gated residual attention for auto rectifying video instance segmentation,, B. Menze, T. Hannan, R. Koner, M. Schubert and, S. Shit, T. Seidl, V. Tresp, M. Bernhard, arXiv preprint arXiv:2305.17096, 2023, , 2023
51. Minvis: A minimal video instance segmentation framework without video-based training, A. Anandkumar, D.-A. Huang, Z. Yu and, Advances in Neural Information Processing Systems, vol. 35, pp. 31 265–31 277, , 2022
52. Classifying, segmenting, and tracking object instances in video with mask propagation, in, L. Torresani, G. Bertasius and, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9739–9748, , 2020
53. Sipmask: Spatial information preservation for fast image and video instance segmentation, in, F. S. Khan, J. Cao, Y. Pang and, R. M. Anwer, H. Cholakkal, L. Shao, Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK,, Proceedings, Part XIV 16. Springer, 2020, pp. 1–18, , 2020
54. Visolo: Grid-based space-time aggregation for efficient online video instance segmentation in, S. J. Kim, Y. Park, H. Kim, S. H. Han, M.-J. Kim and, S. Hwang, S. W. Oh, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2896–2905, , 2022
55. Mask dino: Towards a unified transformer-based framework for object detection and segmentation,, F. Li, L. Zhang, H. Xu, H.-Y. Shum, H. Zhang, S. Liu, L. M. Ni and, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 3041–3050, , 2023
56. Rethinking space-time networks with improved memory coverage for efficient video object segmentation,, H. K. Cheng, C.-K. Tang, Y.-W. Tai and, Advances in Neural Information Processing Systems, vol. 34, pp. 11 781–11 794, , 2021
57. Video object segmentation with adaptive feature bank and uncertain-region refinement in M. Ranzato R., X. Li, Y. Liang, N. Jafari and, J. Chen, M., Advances in Neural Information Processing Systems H. Larochelle Hadsell F. Balcan and H. Lin Eds. vol. 33. Curran Associates, Inc., pp. 3430–3441. [Online]. Available: https://proceedings. neurips. cc/paper/2020/file/ 234833147b97bb6aed53a8f4f1c7a7d8-Paper. pdf, , 2020