RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Trajectory-guided Referred Motion Segmentation = 궤적 특징을 이용한 언어 기반 비디오 내 동작 분할

    한글로보기

    https://www.riss.kr/link?id=T17451111

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Understanding and grounding motion semantics is a fundamental challenge in Referring Video Object Segmentation (RVOS). While recent methods achieve strong performance by integrating spatio-temporal visual features with language, they remain heavily biased toward static appearance cues, often failing when target objects are distinguishable only by motion. In this work, we identify this limitation as the Appearance–Motion Gap and introduce MeViS-D, a diagnostic benchmark that explicitly decouples appearance- and motion-discriminative scenarios. Our analysis reveals a systemic deficiency of existing state-of-the-art models in motion-centric understanding.

    To address this gap, we propose a motion-centric representation based on object trajectories. Rather than relying on concatenated visual features, we leverage the geometric property that trajectories of an object lie in a low-dimensional subspace. By estimating this subspace via truncated SVD and discarding spatial topology, we derive a motion representation that captures pure temporal dynamics while being robust to tracking noise and appearance variations. We further integrate this representation with visual features through a Mixture-of-Experts framework, enabling adaptive, sample-specific balancing between motion and appearance cues.

    Extensive experiments on the MeViS-D benchmark demonstrate that our method establishes a new state-of-the-art on the Motion-Discriminative subset, significantly outperforming appearance-biased baselines. Qualitative analysis further confirms that our model successfully isolates targets from look-alike distractors based solely on temporal patterns, effectively resolving ambiguities where conventional methods fail. These results validate that our structure-agnostic motion representation bridges the Appearance–Motion Gap, offering a robust direction for dynamics-aware video understanding.
    번역하기

    Understanding and grounding motion semantics is a fundamental challenge in Referring Video Object Segmentation (RVOS). While recent methods achieve strong performance by integrating spatio-temporal visual features with language, they remain heavily bi...

    Understanding and grounding motion semantics is a fundamental challenge in Referring Video Object Segmentation (RVOS). While recent methods achieve strong performance by integrating spatio-temporal visual features with language, they remain heavily biased toward static appearance cues, often failing when target objects are distinguishable only by motion. In this work, we identify this limitation as the Appearance–Motion Gap and introduce MeViS-D, a diagnostic benchmark that explicitly decouples appearance- and motion-discriminative scenarios. Our analysis reveals a systemic deficiency of existing state-of-the-art models in motion-centric understanding.

    To address this gap, we propose a motion-centric representation based on object trajectories. Rather than relying on concatenated visual features, we leverage the geometric property that trajectories of an object lie in a low-dimensional subspace. By estimating this subspace via truncated SVD and discarding spatial topology, we derive a motion representation that captures pure temporal dynamics while being robust to tracking noise and appearance variations. We further integrate this representation with visual features through a Mixture-of-Experts framework, enabling adaptive, sample-specific balancing between motion and appearance cues.

    Extensive experiments on the MeViS-D benchmark demonstrate that our method establishes a new state-of-the-art on the Motion-Discriminative subset, significantly outperforming appearance-biased baselines. Qualitative analysis further confirms that our model successfully isolates targets from look-alike distractors based solely on temporal patterns, effectively resolving ambiguities where conventional methods fail. These results validate that our structure-agnostic motion representation bridges the Appearance–Motion Gap, offering a robust direction for dynamics-aware video understanding.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    언어 기반 비디오 내 동작 분할은 언어로 지시된 객체를 비디오 전반에 걸쳐 정확히 분할하는 과제로, 동작 의미를 이해하고 시각 정보와 정합하는 능력이 핵심적인 과제이다. 최근 방법들은 시공간적 시각 특징과 언어 정보를 결합하고 있으나, 여전히 정적 외형 단서에 과도하게 의존하여 동작이 주된 식별 요소인 상황에서 성능이 크게 저하된다. 본 논문은 이를 '외형–동작 갭'으로 정의하고, 외형과 동작 판별력을 분리 진단하는 MeViS-D 벤치마크를 제안하여 최신 모델들이 동작 중심 추론에 구조적인 취약성을 지니고 있음을 보인다.

    이를 해결하기 위해, 본 논문은 객체 궤적에 기반한 동작 중심 표현을 제안한다. 객체 궤적이 저차원 부분공간에 놓인다는 기하학적 성질을 활용하여, 절단된 특이값 분해를 통해 공간적 분포를 제거하고 순수한 시간적 동역학을 포착하는 표현을 도출한다. 또한 Mixture-of-Experts 구조를 통해 동작 표현과 시각 특징을 예제별로 적응적으로 결합한다.

    제안된 방법은 MeViS-D 벤치마크의 동작 판별 평가에서 최고 성능을 달성하며, 정성적 평가를 통해 외형적 단서가 모호한 상황에서도 순수한 시간적 패턴에 기반하여 대상을 정확히 분할할 수 있음을 입증한다. 본 연구는 외형 정보에 편향된 기존 비디오 객체 분할의 한계를 극복하고, 객체 동역학을 비디오 이해의 핵심 요소로 다루는 균형 잡힌 연구의 방향성을 제시한다.
    번역하기

    언어 기반 비디오 내 동작 분할은 언어로 지시된 객체를 비디오 전반에 걸쳐 정확히 분할하는 과제로, 동작 의미를 이해하고 시각 정보와 정합하는 능력이 핵심적인 과제이다. 최근 방법들...

    언어 기반 비디오 내 동작 분할은 언어로 지시된 객체를 비디오 전반에 걸쳐 정확히 분할하는 과제로, 동작 의미를 이해하고 시각 정보와 정합하는 능력이 핵심적인 과제이다. 최근 방법들은 시공간적 시각 특징과 언어 정보를 결합하고 있으나, 여전히 정적 외형 단서에 과도하게 의존하여 동작이 주된 식별 요소인 상황에서 성능이 크게 저하된다. 본 논문은 이를 '외형–동작 갭'으로 정의하고, 외형과 동작 판별력을 분리 진단하는 MeViS-D 벤치마크를 제안하여 최신 모델들이 동작 중심 추론에 구조적인 취약성을 지니고 있음을 보인다.

    이를 해결하기 위해, 본 논문은 객체 궤적에 기반한 동작 중심 표현을 제안한다. 객체 궤적이 저차원 부분공간에 놓인다는 기하학적 성질을 활용하여, 절단된 특이값 분해를 통해 공간적 분포를 제거하고 순수한 시간적 동역학을 포착하는 표현을 도출한다. 또한 Mixture-of-Experts 구조를 통해 동작 표현과 시각 특징을 예제별로 적응적으로 결합한다.

    제안된 방법은 MeViS-D 벤치마크의 동작 판별 평가에서 최고 성능을 달성하며, 정성적 평가를 통해 외형적 단서가 모호한 상황에서도 순수한 시간적 패턴에 기반하여 대상을 정확히 분할할 수 있음을 입증한다. 본 연구는 외형 정보에 편향된 기존 비디오 객체 분할의 한계를 극복하고, 객체 동역학을 비디오 이해의 핵심 요소로 다루는 균형 잡힌 연구의 방향성을 제시한다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents iii
    • 1 Introduction 1
    • 2 Related Works 6
    • Abstract i
    • Contents iii
    • 1 Introduction 1
    • 2 Related Works 6
    • 2.1 Referring Video Object Segmentation 6
    • 2.2 Trajectories as Motion Signal 7
    • 2.3 Trajectory Subspace Clustering 8
    • 3 Method 10
    • 3.1 Problem Formulation 10
    • 3.2 Proposed Method 10
    • 3.2.1 Candidate Object Extraction 10
    • 3.2.2 Motion-Centric Subspace Representation 11
    • 3.2.3 Visual Feature Integration 13
    • 3.2.4 Motion Classification 14
    • 4 Experiments 18
    • 4.1 Experimental Settings 18
    • 4.1.1 Standard Evaluation 18
    • 4.1.2 The MeViS-D Benchmark: A Diagnostic Evaluation 19
    • 4.2 Results 25
    • 4.2.1 Quantitative Results 25
    • 4.2.2 Qualitative Results 28
    • 4.2.3 Efficacy of Structure-Agnostic Representation 29
    • 4.2.4 Efficacy of Visual-Motion Fusion Strategies 35
    • 5 Conclusion 37
    • Abstract (In Korean) 43
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼