Monitoring GPUs at Scale for AI - ML and HPC Clusters

Offered By: CNCF [Cloud Native Computing Foundation] via YouTube

Course Description

Overview

Save Big on Coursera Plus. 7,000+ courses at $160 off. Limited Time Only!

Explore a comprehensive conference talk on monitoring GPU clusters for AI/ML and HPC workloads at scale. Learn how NVIDIA addresses the monitoring needs of various user personas, including AI/ML researchers, operations teams, and stakeholders. Discover the combination of open-source tools used to meet diverse requirements and gain insights into deployment, maintenance, security, and scalability challenges encountered when monitoring GPU data. Understand how NVIDIA overcame these obstacles to create an effective monitoring solution for large GPU Kubernetes clusters running deep learning training workloads.

Syllabus

Monitoring GPUs at Scale for AI/ML and HPC Clusters - Bharti L Agrawal, NVIDIA

Taught by

CNCF [Cloud Native Computing Foundation]

Related Courses

Introduction to Operations Management
Wharton School of the University of Pennsylvania via Coursera Master Control in Supply Chain Management and Logistics
Chalmers University of Technology via edX Supply Chains for Manufacturing: Capacity Analytics
Massachusetts Institute of Technology via edX On Premises Capacity Upgrade and Monitoring with Google Cloud's Apigee API Platform
Google Cloud via Coursera Operations and Supply Chain Management
Indian Institute of Technology Madras via Swayam