Pecan: Cost-Efficient ML Data Preprocessing with Automatic Transformation Ordering and Hybrid Placement
Offered By: USENIX via YouTube
Course Description
Overview
Explore a conference talk from USENIX ATC '24 that introduces Pecan, an innovative ML data preprocessing service designed to optimize costs and efficiency in machine learning workflows. Delve into the challenges of input data preprocessing as a common bottleneck in ML jobs, and discover how Pecan addresses these issues through two key approaches. Learn about the dynamic scheduling of data preprocessing workers on ML accelerator host resources and the automatic reordering of transformations to increase worker throughput. Examine how Pecan's techniques can significantly reduce preprocessing costs by an average of 87% and total training costs by up to 60% compared to existing methods. Gain insights into the importance of balancing commutativity and throughput in ML data pipelines while maintaining high model accuracy. This 21-minute presentation by researchers from ETH Zurich and Google offers valuable knowledge for ML practitioners and researchers looking to enhance the cost-efficiency of their data preprocessing workflows.
Syllabus
USENIX ATC '24 - Pecan: Cost-Efficient ML Data Preprocessing with Automatic Transformation...
Taught by
USENIX
Related Courses
Production Machine Learning SystemsGoogle Cloud via Coursera Deep Learning
Kaggle via YouTube All About AI Accelerators - GPU, TPU, Dataflow, Near-Memory, Optical, Neuromorphic & More
Yannic Kilcher via YouTube Machine Learning with JAX - From Hero to HeroPro+
Aleksa Gordić - The AI Epiphany via YouTube PyTorch NLP Model Training and Fine-Tuning on Colab TPU Multi-GPU with Accelerate
1littlecoder via YouTube