Building the Petcare Data Platform with Delta Lake and Spark ETL Pipeline
Offered By: Databricks via YouTube
Course Description
Overview
Explore the development of Mars Petcare's cloud-based Data Lake solution, the Petcare Data Platform, in this 27-minute conference talk. Learn how the Kinship Data & Analytics division leveraged Microsoft Azure, Delta Lake, and Databricks to create 'Kyte', a custom Spark ETL pipeline tool. Discover the advantages of migrating from Azure Data Factory to a Spark-heavy ETL design and Delta Lake-driven platform. Gain insights into using Delta Lake for ETL configurations and the creation of a bespoke UI for monitoring and scheduling Spark pipelines. Understand the benefits of this approach in supporting Mars Petcare's mission of making a better world for pets, including how Delta Lake is utilized to expose data to Data Scientists and the advantages of a Databricks & Spark ETL solution over Azure Data Factory.
Syllabus
Introduction
Our Data Platform
Our ETL Framework
Schema Revolution
Git Integration
Benefits of Delta Lake
Deployment
Taught by
Databricks
Related Courses
Building Cloud Apps with Microsoft Azure - Part 1 (self-paced)Microsoft via edX Building Cloud Apps with Microsoft Azure - Part 3
Microsoft via edX DEV202.2x: Building Cloud Apps with Microsoft Azure – Part 2
Microsoft via edX Architecting Microsoft Azure Solutions
Microsoft via edX Implementing Predictive Analytics with Spark in Azure HDInsight
Microsoft via edX