RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    A Study on AI-Driven Acoustic Feature Design for Urban Noise Classification

    한글로보기

    https://www.riss.kr/link?id=T17380965

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    A Study on AI-Driven Acoustic Feature Design for Urban Noise Classification SHENYU Department of Mechanical and Automotive Engineering University of Ulsan With the rapid growth of smart-city initiatives, there is a pressing demand for reliable urban acoustic sensing. However, existing resources often lack scene coverage and acoustic diversity, common augmentation practices may introduce artifacts due to insufficient acoustic constraints, and deep neural networks still face challenges in interpretability and deployment robustness. To address these issues, this dissertation follows a knowledge-driven route that tightly couples acoustic domain expertise with modern AI along three coordinated threads: data, modeling, and augmentation with screening. First, a city-oriented dataset, UN15, is constructed from open community audio resources. Fifteen frequently occurring urban sound events are curated via standardized procedures covering duration normalization, dynamic-range control, unified sampling settings, dereverberation/silence handling, and train/validation/test partitioning with class-balance analysis. Experiments indicate that UN15 offers stronger acoustic diversity and application realism than general-purpose benchmarks, serving as an evaluation bed for downstream tasks. Second, a residual network with time – frequency attention, ResNet-TF, is developed. A lightweight attention module is integrated to guide the backbone toward acoustically meaningful transient and stationary structures; cross-domain interactions between time-domain envelopes and frequency-domain spectral shapes are explicitly modeled to capture both impulsive and continuous patterns. Comprehensive experiments and ablations on UN15 show consistent gains over mainstream baselines in overall accuracy and robustness to acquisition variability and noise perturbations. Visual analyses further reveal enhanced and more stable responses to key acoustic cues (e.g., band-limited spectral peaks and modulation sidebands). Third, to mitigate class imbalance and near-neighbor confusions, a boundary-safe audio augmentation strategy anchored to target acoustic feature sets is proposed, together with model-in-the-loop screening. Candidate augmented samples are generated and evaluated in an interpretable feature space (loudness, sharpness, spectral centroid, band-energy ratios, modulation energy, etc.), and jointly filtered by feature-boundary distances and model uncertainty, removing out-of-bounds or confusion-prone instances. Validation on vehicle-sound data demonstrates improved minority-class recognition and cleaner decision boundaries without sacrificing overall performance. Finally, deployment-oriented evaluations are conducted, including computational cost, latency sensitivity, and out-of-domain generalization, complemented by interpretation protocols and error-case diagnostics. The findings demonstrate that systematically embedding acoustic knowledge into data design, model attention, and augmentation/screening policies yields superior trade-offs among accuracy, robustness, and interpretability, significantly enhancing real-world deployability. The dataset, network design, and augmentation-with-screening paradigm presented herein provide reusable baselines and guidelines for engineering-grade urban noise classification. ACKNOWLEGEMENTS At the completion of this dissertation, I would like to express my heartfelt gratitude to all those who have supported and helped me during my doctoral studies. Without their generous assistance and encouragement, this work would not have been possible. First and foremost, I would like to express my deepest appreciation to my advisor, Prof. Chang-Myung Lee, for his invaluable guidance, encouragement, and patience throughout my six years of study and research at the Vibration & Noise Laboratory, University of Ulsan. His rigorous academic attitude, profound expertise, and insightful suggestions have been a constant source of inspiration and motivation for me. I am also sincerely grateful to the members of my dissertation committee for their constructive comments and guidance. In addition, I would like to thank all the professors of the School of Mechanical Engineering who have taught me, for their dedicated instruction that has laid a solid foundation for my academic research. My special thanks also go to the members of the Vibration & Noise Laboratory. I am deeply grateful to Huanyu Dong for his valuable advice and guidance in both research and daily life. I would also like to thank Zhenhua Xu, Min Chen, and Hang Su for their generous sharing of experience and insightful suggestions. I am further indebted to my junior colleague Bo Dong, whose assistance in my experiments and research has been of great help. I would also like to extend my sincere gratitude to my friend and collaborator Ge Cao, with whom I have had many fruitful discussions on research topics, as well as to Yue Teng, for the enjoyable moments we shared outside of work, which brought me relaxation and joy during my doctoral journey. Finally, my deepest gratitude goes to my mother, Yuying Xiao, for her endless love, understanding, encouragement, and patience throughout my long years of study. Her unwavering support has been the greatest driving force behind my academic pursuit.
    번역하기

    A Study on AI-Driven Acoustic Feature Design for Urban Noise Classification SHENYU Department of Mechanical and Automotive Engineering University of Ulsan With the rapid growth of smart-city initiatives, there is a pressing demand for reliable urban a...

    A Study on AI-Driven Acoustic Feature Design for Urban Noise Classification SHENYU Department of Mechanical and Automotive Engineering University of Ulsan With the rapid growth of smart-city initiatives, there is a pressing demand for reliable urban acoustic sensing. However, existing resources often lack scene coverage and acoustic diversity, common augmentation practices may introduce artifacts due to insufficient acoustic constraints, and deep neural networks still face challenges in interpretability and deployment robustness. To address these issues, this dissertation follows a knowledge-driven route that tightly couples acoustic domain expertise with modern AI along three coordinated threads: data, modeling, and augmentation with screening. First, a city-oriented dataset, UN15, is constructed from open community audio resources. Fifteen frequently occurring urban sound events are curated via standardized procedures covering duration normalization, dynamic-range control, unified sampling settings, dereverberation/silence handling, and train/validation/test partitioning with class-balance analysis. Experiments indicate that UN15 offers stronger acoustic diversity and application realism than general-purpose benchmarks, serving as an evaluation bed for downstream tasks. Second, a residual network with time – frequency attention, ResNet-TF, is developed. A lightweight attention module is integrated to guide the backbone toward acoustically meaningful transient and stationary structures; cross-domain interactions between time-domain envelopes and frequency-domain spectral shapes are explicitly modeled to capture both impulsive and continuous patterns. Comprehensive experiments and ablations on UN15 show consistent gains over mainstream baselines in overall accuracy and robustness to acquisition variability and noise perturbations. Visual analyses further reveal enhanced and more stable responses to key acoustic cues (e.g., band-limited spectral peaks and modulation sidebands). Third, to mitigate class imbalance and near-neighbor confusions, a boundary-safe audio augmentation strategy anchored to target acoustic feature sets is proposed, together with model-in-the-loop screening. Candidate augmented samples are generated and evaluated in an interpretable feature space (loudness, sharpness, spectral centroid, band-energy ratios, modulation energy, etc.), and jointly filtered by feature-boundary distances and model uncertainty, removing out-of-bounds or confusion-prone instances. Validation on vehicle-sound data demonstrates improved minority-class recognition and cleaner decision boundaries without sacrificing overall performance. Finally, deployment-oriented evaluations are conducted, including computational cost, latency sensitivity, and out-of-domain generalization, complemented by interpretation protocols and error-case diagnostics. The findings demonstrate that systematically embedding acoustic knowledge into data design, model attention, and augmentation/screening policies yields superior trade-offs among accuracy, robustness, and interpretability, significantly enhancing real-world deployability. The dataset, network design, and augmentation-with-screening paradigm presented herein provide reusable baselines and guidelines for engineering-grade urban noise classification. ACKNOWLEGEMENTS At the completion of this dissertation, I would like to express my heartfelt gratitude to all those who have supported and helped me during my doctoral studies. Without their generous assistance and encouragement, this work would not have been possible. First and foremost, I would like to express my deepest appreciation to my advisor, Prof. Chang-Myung Lee, for his invaluable guidance, encouragement, and patience throughout my six years of study and research at the Vibration & Noise Laboratory, University of Ulsan. His rigorous academic attitude, profound expertise, and insightful suggestions have been a constant source of inspiration and motivation for me. I am also sincerely grateful to the members of my dissertation committee for their constructive comments and guidance. In addition, I would like to thank all the professors of the School of Mechanical Engineering who have taught me, for their dedicated instruction that has laid a solid foundation for my academic research. My special thanks also go to the members of the Vibration & Noise Laboratory. I am deeply grateful to Huanyu Dong for his valuable advice and guidance in both research and daily life. I would also like to thank Zhenhua Xu, Min Chen, and Hang Su for their generous sharing of experience and insightful suggestions. I am further indebted to my junior colleague Bo Dong, whose assistance in my experiments and research has been of great help. I would also like to extend my sincere gratitude to my friend and collaborator Ge Cao, with whom I have had many fruitful discussions on research topics, as well as to Yue Teng, for the enjoyable moments we shared outside of work, which brought me relaxation and joy during my doctoral journey. Finally, my deepest gratitude goes to my mother, Yuying Xiao, for her endless love, understanding, encouragement, and patience throughout my long years of study. Her unwavering support has been the greatest driving force behind my academic pursuit.

    더보기

    목차 (Table of Contents)

    • ABSTRACT i
    • ACKNOWLEGEMENTS iii
    • TABLE OF CONTENTS iv
    • LIST OF FIGURE vi
    • LIST OF TABLE vii
    • ABSTRACT i
    • ACKNOWLEGEMENTS iii
    • TABLE OF CONTENTS iv
    • LIST OF FIGURE vi
    • LIST OF TABLE vii
    • ABBREVIATIONS viii
    • 1. INTRODUCTION 1
    • 1.1 Research Background & Motivation 1
    • 1.2 Key Challenges 2
    • 1.2.1 Data Representativeness 2
    • 1.2.2 Acoustic Rationality & Interpretability 3
    • 1.3 Contributions 4
    • 1.4 Thesis Organization 6
    • 2. Literature Review 8
    • 2.1 Introduction 8
    • 2.2 Development of Audio Classification 9
    • 2.2.1 Pre-deep Era: Feature Engineering & Classical Classifiers 9
    • 2.2.2 Deep-Learning Phase 10
    • 2.3 Environmental Sound Datasets 14
    • 2.3.1 General Benchmarks and Task Definitions 14
    • 2.3.2 Methodological Considerations 16
    • 2.4 Evaluation Metrics System 19
    • 2.5 Data Augmentation and Acoustic Rationality 22
    • 2.5.1 Definition of Data Augmentation 22
    • 2.5.2 General-Purpose Augmentations: Advantages and Limitations 25
    • 2.6 Conclusion 27
    • 3. UN15 Dataset 29
    • 3.1 Introduction 29
    • 3.2 Design Principles and Category Planning 30
    • 3.2.1 Design Principles 30
    • 3.2.2 Category Planning 31
    • 3.3 Experiment and Analysis 33
    • 3.3.1 Data Acquisition 33
    • 3.3.2 Data Annotation 34
    • 3.3.3 Data Preprocessing 35
    • 3.4 Experimental Planning 37
    • 3.5 Summary 38
    • 4. Model Design and Implementation 40
    • 4.1 Introduction 40
    • 4.2 Feature Extraction and Input Representation 40
    • 4.3 Network Architecture and Model Design 42
    • 4.3.1 Network Architecture 42
    • 4.3.2 Temporal-Frequency Attention mechanism 43
    • 4.3.3 Experimental Setup 46
    • 4.4 Results and Analysis 48
    • 4.4.1 Evaluation Metrics and Rationale 48
    • 4.4.2 Overall Experimental Results 49
    • 4.5 Summary 56
    • 5. Boundary-Safe Audio Augmentation 58
    • 5.1 Introduction 58
    • 5.2 Limitations of Conventional Augmentation Under Class Imbalance 60
    • 5.2.1 Principles and Limitations of SpecAugment 61
    • 5.2.2 Limitations of Time-Stretching and Pitch-Shifting 63
    • 5.3 Proposed Boundary-Safe Audio Augmentation 66
    • 5.3.1 Experimental Evaluation and Analysis of SpecAugment 66
    • 5.3.2 Boundary-Safe Audio Augmentation Method 68
    • 5.3.3 Experimental Setup 71
    • 5.4 Results and Analysis 79
    • 6. Conclusion 84
    • 6.1 Summary of Contributions 84
    • 6.2 Future Directions 85
    • Reference 88
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼