The rapid advancement of artificial intelligence and deep learning has positioned Text-to-Audio Generation (TTA) as a significant research focus in fields such as speech synthesis, music composition, and environmental sound effect generation. However,...
The rapid advancement of artificial intelligence and deep learning has positioned Text-to-Audio Generation (TTA) as a significant research focus in fields such as speech synthesis, music composition, and environmental sound effect generation. However, most existing audio generation models are specialized for specific tasks, lacking generalizability and scalability across diverse audio domains. They further face persistent challenges in semantic alignment between text and audio, long-sequence modeling, and controllable generation. To address these challenges, this paper proposes SoundFormer, a cross-modal text-to-audio generation framework integrating BERT, MuseFormer, and SoundStream. By unifying three core modules—text semantic understanding (BERT), symbolic music generation (MuseFormer), and high-fidelity audio synthesis (SoundStream)—SoundFormer can generate various types of audio, including speech, music, and sound effects, from natural language descriptions. This research focuses on enhancing semantic consistency between the input text and the generated audio through hierarchical attention mechanisms and multi-objective loss functions, ensuring high quality in both naturalness and emotional expressiveness. Comparative experiments, ablation studies, and subjective evaluations demonstrate SoundFormer's superior performance in audio quality, semantic alignment, and multi-task generation capability. The primary contributions of this work are fourfold: (1) the proposal and implementation of the SoundFormer model as an end-to-end TTA framework; (2) the design of an innovative cross-modal alignment mechanism that improves semantic consistency between text and audio; (3) the construction of a large-scale multimodal dataset to enhance the diversity of training data; and (4) an exploration of SoundFormer's potential in practical applications such as virtual reality, game audio, and intelligent content creation. Ultimately, this paper provides a unified solution for text-driven audio generation characterized by high flexibility, controllability, and extensibility.