A Study on AI-Driven Acoustic Feature Design for Urban Noise Classification SHENYU Department of Mechanical and Automotive Engineering University of Ulsan With the rapid growth of smart-city initiatives, there is a pressing demand for reliable urban a...
A Study on AI-Driven Acoustic Feature Design for Urban Noise Classification SHENYU Department of Mechanical and Automotive Engineering University of Ulsan With the rapid growth of smart-city initiatives, there is a pressing demand for reliable urban acoustic sensing. However, existing resources often lack scene coverage and acoustic diversity, common augmentation practices may introduce artifacts due to insufficient acoustic constraints, and deep neural networks still face challenges in interpretability and deployment robustness. To address these issues, this dissertation follows a knowledge-driven route that tightly couples acoustic domain expertise with modern AI along three coordinated threads: data, modeling, and augmentation with screening. First, a city-oriented dataset, UN15, is constructed from open community audio resources. Fifteen frequently occurring urban sound events are curated via standardized procedures covering duration normalization, dynamic-range control, unified sampling settings, dereverberation/silence handling, and train/validation/test partitioning with class-balance analysis. Experiments indicate that UN15 offers stronger acoustic diversity and application realism than general-purpose benchmarks, serving as an evaluation bed for downstream tasks. Second, a residual network with time – frequency attention, ResNet-TF, is developed. A lightweight attention module is integrated to guide the backbone toward acoustically meaningful transient and stationary structures; cross-domain interactions between time-domain envelopes and frequency-domain spectral shapes are explicitly modeled to capture both impulsive and continuous patterns. Comprehensive experiments and ablations on UN15 show consistent gains over mainstream baselines in overall accuracy and robustness to acquisition variability and noise perturbations. Visual analyses further reveal enhanced and more stable responses to key acoustic cues (e.g., band-limited spectral peaks and modulation sidebands). Third, to mitigate class imbalance and near-neighbor confusions, a boundary-safe audio augmentation strategy anchored to target acoustic feature sets is proposed, together with model-in-the-loop screening. Candidate augmented samples are generated and evaluated in an interpretable feature space (loudness, sharpness, spectral centroid, band-energy ratios, modulation energy, etc.), and jointly filtered by feature-boundary distances and model uncertainty, removing out-of-bounds or confusion-prone instances. Validation on vehicle-sound data demonstrates improved minority-class recognition and cleaner decision boundaries without sacrificing overall performance. Finally, deployment-oriented evaluations are conducted, including computational cost, latency sensitivity, and out-of-domain generalization, complemented by interpretation protocols and error-case diagnostics. The findings demonstrate that systematically embedding acoustic knowledge into data design, model attention, and augmentation/screening policies yields superior trade-offs among accuracy, robustness, and interpretability, significantly enhancing real-world deployability. The dataset, network design, and augmentation-with-screening paradigm presented herein provide reusable baselines and guidelines for engineering-grade urban noise classification. ACKNOWLEGEMENTS At the completion of this dissertation, I would like to express my heartfelt gratitude to all those who have supported and helped me during my doctoral studies. Without their generous assistance and encouragement, this work would not have been possible. First and foremost, I would like to express my deepest appreciation to my advisor, Prof. Chang-Myung Lee, for his invaluable guidance, encouragement, and patience throughout my six years of study and research at the Vibration & Noise Laboratory, University of Ulsan. His rigorous academic attitude, profound expertise, and insightful suggestions have been a constant source of inspiration and motivation for me. I am also sincerely grateful to the members of my dissertation committee for their constructive comments and guidance. In addition, I would like to thank all the professors of the School of Mechanical Engineering who have taught me, for their dedicated instruction that has laid a solid foundation for my academic research. My special thanks also go to the members of the Vibration & Noise Laboratory. I am deeply grateful to Huanyu Dong for his valuable advice and guidance in both research and daily life. I would also like to thank Zhenhua Xu, Min Chen, and Hang Su for their generous sharing of experience and insightful suggestions. I am further indebted to my junior colleague Bo Dong, whose assistance in my experiments and research has been of great help. I would also like to extend my sincere gratitude to my friend and collaborator Ge Cao, with whom I have had many fruitful discussions on research topics, as well as to Yue Teng, for the enjoyable moments we shared outside of work, which brought me relaxation and joy during my doctoral journey. Finally, my deepest gratitude goes to my mother, Yuying Xiao, for her endless love, understanding, encouragement, and patience throughout my long years of study. Her unwavering support has been the greatest driving force behind my academic pursuit.