CUTLASS: A CUDA C++ Template Library for Accelerating Deep Learning Computations
Offered By: Linux Foundation via YouTube
Course Description
Overview
Explore CUTLASS, an open-source CUDA C++ template library designed to accelerate deep learning computations on NVIDIA GPUs. Dive into the core concepts of GPU computing for machine learning and AI applications, focusing on optimizing linear algebra operations like matrix multiplication and convolutions. Learn how CUTLASS has been instrumental since 2017 in helping developers create high-performance CUDA kernels across various NVIDIA GPU architectures. Gain insights into Tensor Core programming and discover how to leverage CUTLASS's modular abstractions and building blocks to develop custom CUDA C++ kernels that maximize performance for deep learning tasks. Acquire actionable knowledge to push the limits of GPU performance in AI applications like ChatGPT and Github Copilot.
Syllabus
CUTLASS: A CUDA C++ Template Library for Accelerating Deep Learning... Aniket Shivam & Vijay Thakkar
Taught by
Linux Foundation
Tags
Related Courses
Neural Networks for Machine LearningUniversity of Toronto via Coursera 機器學習技法 (Machine Learning Techniques)
National Taiwan University via Coursera Machine Learning Capstone: An Intelligent Application with Deep Learning
University of Washington via Coursera Прикладные задачи анализа данных
Moscow Institute of Physics and Technology via Coursera Leading Ambitious Teaching and Learning
Microsoft via edX