AudioGen- Textually Guided Audio Generation - Paper Explained
Offered By: Aleksa Gordić - The AI Epiphany via YouTube
Course Description
Overview
Dive deep into the world of text-guided audio synthesis with this comprehensive video explanation of the "AudioGen: Textually Guided Audio Generation" paper. Explore the challenges of text-to-audio conversion, compare AudioGen with VQ-GAN and SoundStream, and gain insights into audio representation, LSTM networks, and complex-valued STFTs. Learn about audio language modeling, multi-stream audio inputs, data augmentation techniques, and examine the impressive results of this innovative approach to audio generation.
Syllabus
Intro
Why is text-to-audio hard?
Comparison with VQ-GAN
Comparison with SoundStream
AudioGen overview
Deep dive: audio representation, LSTM
Losses explained
Complex-valued STFTs
Audio Language Modeling
Multi-stream audio inputs
Data and augmentations
Results
Outro
Taught by
Aleksa Gordić - The AI Epiphany
Related Courses
TensorFlow を使った畳み込みニューラルネットワークDeepLearning.AI via Coursera Emotion AI: Facial Key-points Detection
Coursera Project Network via Coursera Transfer Learning for Food Classification
Coursera Project Network via Coursera Facial Expression Classification Using Residual Neural Nets
Coursera Project Network via Coursera Apply Generative Adversarial Networks (GANs)
DeepLearning.AI via Coursera