YoVDO

Bytedance's Parquet Format Optimization for Cost Reduction and Efficiency

Offered By: The ASF via YouTube

Tags

Parquet Courses SQL Courses Data Warehousing Courses Apache Spark Courses Data Storage Courses

Course Description

Overview

Save Big on Coursera Plus. 7,000+ courses at $160 off. Limited Time Only!
Explore Bytedance's innovative approach to cost reduction and efficiency improvement using the Parquet format in this informative conference talk. Discover how Bytedance tackled challenges related to small file proliferation and high data storage costs in their offline data warehouse. Learn about the advanced techniques implemented to optimize Parquet file overwriting, including binary copy methods that bypass redundant operations like codec and decompression. Gain insights into the performance improvements achieved, with efficiency gains of over 10 times compared to traditional overwriting methods. Understand the new SQL syntax introduced to simplify small file merging and column-level TTL operations, enhancing user experience and data management capabilities.

Syllabus

Bytedance Based On The Parquet Format Of Cost Reduction And Efficiency Practice


Taught by

The ASF

Related Courses

Software Engineering for SaaS
University of California, Berkeley via Coursera
Android. Programación de Aplicaciones
Miríadax
Informatik für Ökonomen
University of Zurich via Coursera
Aprendizaje en la Nube, Herramientas web en el Aula
Galileo University via Independent
Desarrollo de Aplicaciones para Android
Galileo University via Independent