Creating Voiceovers with OpenAI's Text-to-Speech and Vision Models
Offered By: Ian Wootten via YouTube
Course Description
Overview
Explore the latest advancements in OpenAI's text-to-speech (TTS) and GPT-4V models in this 15-minute video tutorial. Discover innovative applications of these technologies, including generating image descriptions and creating audio content. Learn how to produce voiceovers for images and videos using a combination of TTS and GPT-4V. Follow along as the presenter demonstrates practical examples and showcases novel ways developers have been utilizing these powerful tools. Gain insights into the potential of AI-driven content creation and enhance your understanding of cutting-edge language and vision models.
Syllabus
Intro
Using TTS to create audio
Using GPT4V to describe images
Using TTS & GPT4V for Video voiceovers
Conclusion
Taught by
Ian Wootten
Related Courses
AudioGen- Textually Guided Audio Generation - Paper ExplainedAleksa Gordić - The AI Epiphany via YouTube MusicLM Generates Music From Text - Paper Breakdown
Valerio Velardo - The Sound of AI via YouTube A Composer's Guide to Creating with Generative Neural Networks
GOTO Conferences via YouTube 21 Recent AI Updates in 23 Minutes
1littlecoder via YouTube Popcorn & Clocks - A Story About Scheduling in the Browser
NDC Conferences via YouTube