Cardinality Estimation in Apache Spark 2.3 (Ron Hu & Zhenhua Wang)
431 views · Published 27 September 2018 · 30:15 · Indexed 21 September 2026
Channel: Databricks · 2018 · Science & Technology
Ron Hu, a Principal Big Data Architect at Huawei Technologies, and Zhenhua Wang, a Research Engineer at Huawei Technologies, explain how Apache Spark 2.2 shipped with a state-of-art cost-based optimization framework that collects and leverages a variety of per-column data statistics (e.g., cardinality, number of distinct values, NULL values, max/min, avg/max length, etc.) to improve the quality of query execution plans. Skewed data distributions are often inherent in many real world applications. In order to deal with skewed distributions effectively, we added equal-height histograms to Apache Spark 2.3. Leveraging reliable statistics and histogram helps Spark make better decisions in picking the most optimal query plan for real world scenarios. Learn more here: https://databricks.com/session/cardinality-estimation-through-histogram-in-apache-spark-2-3 Article you might like: https://databricks.com/session/deep-learning-for-natural-language-processing-using-apache-spark-tensorflow About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business. Read more here: https://databricks.com/product/unified-data-analytics-platform Connect with us: Website: https://databricks.com Facebook: https://www.facebook.com/databricksinc Twitter: https://twitter.com/databricks LinkedIn: https://www.linkedin.com/company/databricks Instagram: https://www.instagram.com/databricksinc/ Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-named-leader-by-gartner
More from this channel
-
28:55
Scaling Up Data Science Applications with Kexin Xie and Yacov Salomon
-
31:03
Extending Spark Machine Learning: Adding Your Own Algorithms and Tools
-
16:23
SparkOscope: Enabling Apache Spark Optimization through Cross Stack Monitoring - Yiannis Gkoufas
-
19:53
Women in Big Data Lunch
-
33:31
Improving Traffic Prediction Using Weather Data - Ramya Raghavendra
-
30:28
Efficiently Triaging CI Pipelines with Apache Spark (Ivan Jibaja)
-
31:45
Bringing an AI Ecosystem to the Domain Expert and Enterprise AI Developer (Frederick Reiss)
-
29:47
Time Series Anomaly Detection in Plaintext Using Apache Spark with Jerry Schirmer (SparkCognition)