V-JEPA: Revisiting Feature Prediction for Learning Visual Representations from Video
Offered By: Yannic Kilcher via YouTube
Course Description
Overview
Explore an in-depth explanation of V-JEPA (Video Joint Embedding Predictive Architecture), a novel method for unsupervised representation learning from video data. Delve into the predictive feature principle, the original JEPA architecture, and the V-JEPA concept and architecture. Examine experimental results and qualitative evaluation through decoding. Learn how this approach, developed by Meta AI researchers, achieves impressive performance on both motion and appearance-based tasks using only latent representation prediction as an objective function. Gain insights into the potential of this technique for advancing unsupervised learning in computer vision and its implications for future AI developments.
Syllabus
- Intro
- Predictive Feature Principle
- Weights & Biases course on Structured LLM Outputs
- The original JEPA architecture
- V-JEPA Concept
- V-JEPA Architecture
- Experimental Results
- Qualitative Evaluation via Decoding
Taught by
Yannic Kilcher
Related Courses
Machine Learning: Unsupervised LearningBrown University via Udacity Practical Predictive Analytics: Models and Methods
University of Washington via Coursera Поиск структуры в данных
Moscow Institute of Physics and Technology via Coursera Statistical Machine Learning
Carnegie Mellon University via Independent FA17: Machine Learning
Georgia Institute of Technology via edX