Most of pedestrian attribute recognition (PAR) exploit only visual cues, which has high dependency for image conditions and would lead sub-optimal results. In this paper, we adopt the text descriptions as auxiliary information for PAR. For low trainin...
Most of pedestrian attribute recognition (PAR) exploit only visual cues, which has high dependency for image conditions and would lead sub-optimal results. In this paper, we adopt the text descriptions as auxiliary information for PAR. For low training cost but sufficient multi-modal represenstation space, we introduce the mluti-modal prompt learning. Specifically, we construct visual and text prompt which consist of learnable parameters, and it go though the fixed image and text encoders, respectively. To focus effectively on the relationship between different modalities, we introduce the XSA module to fuse the tokens. Our proposed PAR model achieves accuracy of 79.28 and 86.67 of F1-score on PETA dataset. Various experiments are conducted to validate the effectiveness of our model.