Bytedance's Parquet Format Optimization for Cost Reduction and Efficiency
Offered By: The ASF via YouTube
Course Description
Overview
Explore Bytedance's innovative approach to cost reduction and efficiency improvement using the Parquet format in this informative conference talk. Discover how Bytedance tackled challenges related to small file proliferation and high data storage costs in their offline data warehouse. Learn about the advanced techniques implemented to optimize Parquet file overwriting, including binary copy methods that bypass redundant operations like codec and decompression. Gain insights into the performance improvements achieved, with efficiency gains of over 10 times compared to traditional overwriting methods. Understand the new SQL syntax introduced to simplify small file merging and column-level TTL operations, enhancing user experience and data management capabilities.
Syllabus
Bytedance Based On The Parquet Format Of Cost Reduction And Efficiency Practice
Taught by
The ASF
Related Courses
CS115x: Advanced Apache Spark for Data Science and Data EngineeringUniversity of California, Berkeley via edX Big Data Analytics
University of Adelaide via edX Big Data Essentials: HDFS, MapReduce and Spark RDD
Yandex via Coursera Big Data Analysis: Hive, Spark SQL, DataFrames and GraphFrames
Yandex via Coursera Introduction to Apache Spark and AWS
University of London International Programmes via Coursera