Understanding and grounding motion semantics is a fundamental challenge in Referring Video Object Segmentation (RVOS). While recent methods achieve strong performance by integrating spatio-temporal visual features with language, they remain heavily bi...
Understanding and grounding motion semantics is a fundamental challenge in Referring Video Object Segmentation (RVOS). While recent methods achieve strong performance by integrating spatio-temporal visual features with language, they remain heavily biased toward static appearance cues, often failing when target objects are distinguishable only by motion. In this work, we identify this limitation as the Appearance–Motion Gap and introduce MeViS-D, a diagnostic benchmark that explicitly decouples appearance- and motion-discriminative scenarios. Our analysis reveals a systemic deficiency of existing state-of-the-art models in motion-centric understanding.
To address this gap, we propose a motion-centric representation based on object trajectories. Rather than relying on concatenated visual features, we leverage the geometric property that trajectories of an object lie in a low-dimensional subspace. By estimating this subspace via truncated SVD and discarding spatial topology, we derive a motion representation that captures pure temporal dynamics while being robust to tracking noise and appearance variations. We further integrate this representation with visual features through a Mixture-of-Experts framework, enabling adaptive, sample-specific balancing between motion and appearance cues.
Extensive experiments on the MeViS-D benchmark demonstrate that our method establishes a new state-of-the-art on the Motion-Discriminative subset, significantly outperforming appearance-biased baselines. Qualitative analysis further confirms that our model successfully isolates targets from look-alike distractors based solely on temporal patterns, effectively resolving ambiguities where conventional methods fail. These results validate that our structure-agnostic motion representation bridges the Appearance–Motion Gap, offering a robust direction for dynamics-aware video understanding.