Lazy Join Optimizations Without Upfront Statistics - Matteo Interlandi
418 views · Published 8 June 2017 · 30:50 · Indexed 29 September 2026
Channel: Databricks · 2017 · Science & Technology
Modern Data-Intensive Scalable Computing (DISC) systems such as Apache Spark do not support sophisticated cost-based query optimizers because they are specifically designed to process data that resides in external storage systems (e.g. HDFS), or they lack the necessary data statistics. Consequently, many crucial optimizations, such as join order and plan selection, are presently out-of-scope in these DISC system optimizers. Yet, join order is one of the most important decisions a cost-optimizer can make because wrong orders can result in a query response time that can become more than an order-of-magnitude slower compared to the better order. About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business. Read more here: https://databricks.com/product/unified-data-analytics-platform Connect with us: Website: https://databricks.com Facebook: https://www.facebook.com/databricksinc Twitter: https://twitter.com/databricks LinkedIn: https://www.linkedin.com/company/databricks Instagram: https://www.instagram.com/databricksinc/ Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-named-leader-by-gartner
More from this channel
-
10:06
Big Data Meets Learning Science
-
30:18
Yelp Ad Targeting at Scale with Apache Spark - Inaz Alaei-Novin and Joe Malicki
-
37:42
Scaling Data Science Capabilities with Apache Spark at Stitch Fix - Derek Bennett
-
31:04
Debugging Big Data Analytics in Apache Spark with BigDebug Matteo Interlandi and Muhammad Ali Gulzar
-
28:55
Scaling Up Data Science Applications with Kexin Xie and Yacov Salomon
-
31:03
Extending Spark Machine Learning: Adding Your Own Algorithms and Tools
-
23:09
The Key to Machine Learning is Prepping the Right Data - Jean Georges Perrin
-
30:16
Best Practices for Using Alluxio with Apache Spark - Cheng Chang & Haoyuan Li