SLA-Aware Machine Learning Inference Serving on Serverless Computing Platforms
Offered By: MLOps World: Machine Learning in Production via YouTube
Course Description
Overview
Explore a conference talk on SLA-aware machine learning inference serving on serverless computing platforms. Delve into the challenges of serving machine learning inference workloads in production environments and the complexities of meeting SLA requirements while optimizing infrastructure costs. Learn about MLProxy, an adaptive reverse proxy designed to support efficient machine learning serving workloads on serverless systems. Discover how MLProxy utilizes adaptive batching to ensure SLA compliance and optimize serverless costs. Examine the results of rigorous experiments conducted on Knative, demonstrating MLProxy's ability to significantly reduce serverless deployment costs and SLA violations across various model serving frameworks.
Syllabus
SLA Aware Machine Learning Inference Serving on Serverless Computing Platforms
Taught by
MLOps World: Machine Learning in Production
Related Courses
Introduction to Cloud Infrastructure TechnologiesLinux Foundation via edX Cloud Computing
Indian Institute of Technology, Kharagpur via Swayam Elastic Cloud Infrastructure: Containers and Services en Español
Google Cloud via Coursera Kyma – A Flexible Way to Connect and Extend Applications
SAP Learning Modernize Infrastructure and Applications with Google Cloud
Google Cloud via Coursera