RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Video Scene Graph Generation via Temporal Relation Change Modeling = 시간적 관계 변화 모델링을 통한 비디오 장면 그래프 생성 연구

    한글로보기

    https://www.riss.kr/link?id=T17570779

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    For machine intelligence to extract high-level semantic information from video data and apply it to real-world interaction and problem-solving, two capabilities are essential: the ability to identify the entities present in a scene, and the ability to recognize the dynamic interactions that unfold among them over time. Video Scene Graph Generation (VidSGG) provides a principled framework for such structured semantic extraction, representing a video sequence as a temporally evolving graph through the systematic parsing of ⟨subject, predicate, object⟩ triplets. In this thesis, we analyze the limitations of existing VidSGG approaches and propose a novel framework to address them.
    Accurately modeling the evolution of inter-object relations as a video sequence progresses remains a fundamental challenge. This difficulty arises because the temporal aggregation strategies of existing methods lack sufficient mechanisms to capture the high variability of relational dynamics. Consequently, existing models suffer from Relational Semantic Inertia — an inability to promptly detect transition boundaries when inter-object relations shift.
    To overcome this problem, we propose the Relation-Transition Aware Transformer (RTAformer), a framework that explicitly models changes in relation representations through three complementary modules and incorporates transition information during relation prediction. First, Spatio-Semantic Instance Alignment (SSIA) applies contrastive regularization to instance-pair representations across adjacent frames, forming temporally consistent representations for the same subject–object pair so that subsequent relation-query differences reflect genuine semantic transitions. Second, Disentangled Temporal Aggregation (DTA) separates the temporal refinement of instance and relation information into two independent streams: the instance stream performs persistence-oriented aggregation, while the relation stream is dedicated to transition-sensitive processing. Third, Transition-Aware Delta Injection (TADI) decomposes the adjacent-frame relation-query displacement into a direction component and a magnitude component, modulates the resulting transition signal, and injects it into the relation prediction query — enabling the model to respond promptly at relational transition boundaries.
    The proposed method is validated on the Action Genome benchmark through quantitative and qualitative experiments, demonstrating superior performance over existing VidSGG methods. In particular, a dedicated evaluation on a constructed relation-transition subset shows that RTAformer improves transition-boundary prediction over representative existing methods, supporting the effectiveness of transition-aware temporal modeling in mitigating relational semantic inertia. The proposed framework is expected to advance machine intelligence by enabling more accurate semantic extraction from video data, with broad applicability to embodied AI, grounded reasoning, and holistic video understanding.
    번역하기

    For machine intelligence to extract high-level semantic information from video data and apply it to real-world interaction and problem-solving, two capabilities are essential: the ability to identify the entities present in a scene, and the ability to...

    For machine intelligence to extract high-level semantic information from video data and apply it to real-world interaction and problem-solving, two capabilities are essential: the ability to identify the entities present in a scene, and the ability to recognize the dynamic interactions that unfold among them over time. Video Scene Graph Generation (VidSGG) provides a principled framework for such structured semantic extraction, representing a video sequence as a temporally evolving graph through the systematic parsing of ⟨subject, predicate, object⟩ triplets. In this thesis, we analyze the limitations of existing VidSGG approaches and propose a novel framework to address them.
    Accurately modeling the evolution of inter-object relations as a video sequence progresses remains a fundamental challenge. This difficulty arises because the temporal aggregation strategies of existing methods lack sufficient mechanisms to capture the high variability of relational dynamics. Consequently, existing models suffer from Relational Semantic Inertia — an inability to promptly detect transition boundaries when inter-object relations shift.
    To overcome this problem, we propose the Relation-Transition Aware Transformer (RTAformer), a framework that explicitly models changes in relation representations through three complementary modules and incorporates transition information during relation prediction. First, Spatio-Semantic Instance Alignment (SSIA) applies contrastive regularization to instance-pair representations across adjacent frames, forming temporally consistent representations for the same subject–object pair so that subsequent relation-query differences reflect genuine semantic transitions. Second, Disentangled Temporal Aggregation (DTA) separates the temporal refinement of instance and relation information into two independent streams: the instance stream performs persistence-oriented aggregation, while the relation stream is dedicated to transition-sensitive processing. Third, Transition-Aware Delta Injection (TADI) decomposes the adjacent-frame relation-query displacement into a direction component and a magnitude component, modulates the resulting transition signal, and injects it into the relation prediction query — enabling the model to respond promptly at relational transition boundaries.
    The proposed method is validated on the Action Genome benchmark through quantitative and qualitative experiments, demonstrating superior performance over existing VidSGG methods. In particular, a dedicated evaluation on a constructed relation-transition subset shows that RTAformer improves transition-boundary prediction over representative existing methods, supporting the effectiveness of transition-aware temporal modeling in mitigating relational semantic inertia. The proposed framework is expected to advance machine intelligence by enabling more accurate semantic extraction from video data, with broad applicability to embodied AI, grounded reasoning, and holistic video understanding.

    더보기

    목차 (Table of Contents)

    • 1 Introduction 1
    • 1.1 Motivation 1
    • 1.2 Related Work 6
    • 1.2.1 Video Scene Graph Generation 6
    • 1.2.2 Temporal Relational Dynamics and Transitions 6
    • 1 Introduction 1
    • 1.1 Motivation 1
    • 1.2 Related Work 6
    • 1.2.1 Video Scene Graph Generation 6
    • 1.2.2 Temporal Relational Dynamics and Transitions 6
    • 1.2.3 Query Evolution and Temporal Initialization 7
    • 1.3 Contributions 9
    • 1.4 Thesis Outline 10
    • 2 Proposed Video Scene Graph Generation Method 11
    • 2.1 Problem Formulation 11
    • 2.2 Framework Overview 11
    • 2.3 Frame-wise Spatial Triplet Parsing with Consistency Regularization 13
    • 2.3.1 Frame-wise Spatial Triplet Parsing 13
    • 2.3.2 Spatio-Semantic Instance Alignment 15
    • 2.4 Disentangled Temporal Aggregation 17
    • 2.4.1 Persistence-Driven Instance Trajectory Smoothing 19
    • 2.4.2 Transition-Aware Delta Injection (TADI) 21
    • 2.5 Optimization and Training Objectives 25
    • 2.5.1 Decoupled Training Curriculum 25
    • 2.5.2 Bipartite Matching and Stage-wise Objectives 26
    • 2.5.3 One-Stage Inference 28
    • 3 Experimental Results and Discussion 29
    • 3.1 Experimental Setup 29
    • 3.1.1 Benchmark 29
    • 3.1.2 Evaluation Metrics 29
    • 3.1.3 Implementation Details 30
    • 3.2 Comparison with Existing VidSGG Methods 32
    • 3.2.1 Scene Graph Detection Recall Comparison 32
    • 3.2.2 Mean Recall Comparison 33
    • 3.3 Ablation Studies 34
    • 3.3.1 Effect of SSIA and DTA 34
    • 3.3.2 Representation Analysis of SSIA 35
    • 3.3.3 Effect of Transition-Aware Delta Injection 37
    • 3.3.4 Effect of Temporal Aggregation Depth 38
    • 3.4 Analysis of Relational Semantic Inertia 39
    • 3.4.1 Relation-Transition Subset Evaluation 39
    • 3.4.2 Qualitative Scene Graph Comparison 41
    • 4 Conclusion 42
    • References 45
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼