MAST: Global Scheduling of ML Training across Geo-Distributed Datacenters at Hyperscale
Offered By: USENIX via YouTube
Course Description
Overview
Save Big on Coursera Plus. 7,000+ courses at $160 off. Limited Time Only!
Explore a conference talk on MAST, a global scheduler for ML training workloads across geo-distributed datacenters at hyperscale. Learn about the challenges of manual datacenter region selection in public clouds and how MAST addresses these issues in Meta's private cloud. Discover the three key design principles enabling MAST to schedule complex ML training workloads globally: temporal decoupling, scope decoupling, and exhaustive search. Understand how MAST successfully balances load across global regions, reducing the GPU demand-to-supply ratio for high-priority workloads from 2.63 to 0.98 in the most overloaded region. Gain insights into the global-scheduling abstraction provided by MAST and its impact on hardware utilization and profitability.
Syllabus
OSDI '24 - MAST: Global Scheduling of ML Training across Geo-Distributed Datacenters at Hyperscale
Taught by
USENIX
Related Courses
A Hands-On Look at Amazon Q Business ExpertAmazon Web Services via AWS Skill Builder À la découverte des télécommunications
Institut Mines-Télécom via France Université Numerique A Tour of Google Cloud Sustainability
Google via Google Cloud Skills Boost Intel® Telco Cloud Academy
Intel via Coursera Accéder à Internet depuis Lambda dans un VPC (Français) | Accessing the Internet from Lambda in a VPC (French)
Amazon Web Services via AWS Skill Builder