The contemporary fashion industry has undergone a fundamental transformation toward digital platforms and social media, creating unprecedented opportunities alongside substantial challenges in visual product discovery, data-driven trend forecasting, a...
The contemporary fashion industry has undergone a fundamental transformation toward digital platforms and social media, creating unprecedented opportunities alongside substantial challenges in visual product discovery, data-driven trend forecasting, and personalized styling assistance. This dissertation addresses these challenges through an integrated framework of AI-powered systems for fashion data representation and application.
This research comprises three interconnected components that collectively establish intelligent fashion systems for online environments. First, we develop a cross-modal fashion image retrieval system that bridges the significant visual gap between user-generated street photos and professional product catalogs. Leveraging a sigmoid-based contrastive learning framework (SigLIP) with attribute-aware pre-training and weighted contrastive fine-tuning, our approach achieves state-of-the-art performance on Street2Shop and DeepFashion benchmarks, demonstrating top-1 accuracy improvements of 26.3% and 45.5% respectively over existing methods.
Second, we construct the Fashion Visual Instruction (FVI) dataset comprising 5 million hierarchical instruction samples across three complexity levels: Simple (factual classification), Complex (multi-attribute extraction), and Advanced (logical reasoning). By applying Localized Fine-Tuning to Qwen2-VL with bounding box integration, we enable the model to perform precise attribute extraction and contextual inference from unstructured, cluttered social media images. The model demonstrates superior capability in analyzing temporal fashion dynamics, achieving strong performance across quantitative metrics (BERTScore: 0.942, METEOR: 0.614), LLM-assisted evaluation (LAVE: 2.370), and human expert assessment (4.35/5.0).
Third, we integrate these components into an end-to-end multi-modal conversational agent utilizing Retrieval-Augmented Generation (RAG), domain-specific alignment via Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO), and visual grounding through cross-modal retrieval. Our specialized 8B parameter model achieves a 58% human preference rate, surpassing GPT-4o's 40%, while demonstrating superior linguistic fluency (Perplexity: 12.28) and lexical diversity (Dist-3: 0.856). The system successfully transforms quantitative trend insights into accessible, personalized styling advice supported by concrete visual evidence.
The synergistic integration of visual retrieval, trend analysis, and conversational interaction establishes a comprehensive fashion intelligence framework. This work demonstrates that domain-specific adaptation and multi-modal integration enable AI systems to provide practical, user-centric fashion assistance. By combining data-driven insights with engaging interaction, this dissertation presents a pathway to making professional fashion expertise accessible to consumers in the digital era.