Dask-SQL - Empowering Pythonistas for Scalable End-to-End Data Engineering
Offered By: PyCon US via YouTube
Course Description
Overview
Discover how to leverage dask-sql for scalable end-to-end data engineering in this 26-minute PyCon US talk. Learn to overcome the challenges of accessing data trapped in Hive/Spark-based datalakes or complex SQL queries. Explore the capabilities of dask-sql, which enables Python developers to create comprehensive data projects without extensive knowledge of JVM/Hadoop ecosystems. Gain insights into performing SQL data extraction from datalakes and Hive tables using Python and dask-sql. Understand how to refine and utilize extracted data for machine learning, analytics, or transformation workloads with popular PyData tools. Delve into the innovative design of dask-sql, which combines SQL optimization from Apache Calcite, scalable dataframe operations via Dask, and integration with the Hive metastore data catalog. Follow along with a demo and walkthrough of dask-sql, covering topics such as the enterprise data processing pipeline and practical implementation. Access accompanying slides for further reference and study.
Syllabus
Introduction
Enterprise Data Processing Pipeline
DaskSQL
Demo
DaskSQL Walkthrough
Taught by
PyCon US
Related Courses
Introduction to Artificial IntelligenceStanford University via Udacity Natural Language Processing
Columbia University via Coursera Probabilistic Graphical Models 1: Representation
Stanford University via Coursera Computer Vision: The Fundamentals
University of California, Berkeley via Coursera Learning from Data (Introductory Machine Learning course)
California Institute of Technology via Independent