Spark SQL Adaptive Execution Unleashes The Power of Cluster (Carson Wang and Yuanjian Li)
833 views · Published 24 September 2018 · 25:46 · Indexed 26 September 2026
Channel: Databricks · 2018 · Science & Technology
Carson Wang is a big data software engineer at Intel, focusing on developing and improving new big data technologies. Yuanjian is a senior engineer and the lead for distributed computing team at Baidu. In this talk, we will explore Intel and Baidu’s joint efforts to address challenges in large scale and offer an overview of an adaptive execution mode we implemented for Baidu’s Big SQL platform which is based on Spark SQL. At runtime, adaptive execution can change the execution plan to use a better join strategy and handle skewed join automatically. It can also change the number of reducer to better fit the data scale. In general, adaptive execution decreases the effort involved in tuning SQL query parameters and improves the execution performance by choosing a better execution plan and parallelism at runtime. To learn more: https://databricks.com/product/getting-started-guide About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business. Read more here: https://databricks.com/product/unified-data-analytics-platform Connect with us: Website: https://databricks.com Facebook: https://www.facebook.com/databricksinc Twitter: https://twitter.com/databricks LinkedIn: https://www.linkedin.com/company/databricks Instagram: https://www.instagram.com/databricksinc/ Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-named-leader-by-gartner
More from this channel
-
10:06
Big Data Meets Learning Science
-
28:55
Scaling Up Data Science Applications with Kexin Xie and Yacov Salomon
-
31:03
Extending Spark Machine Learning: Adding Your Own Algorithms and Tools
-
23:09
The Key to Machine Learning is Prepping the Right Data - Jean Georges Perrin
-
16:23
SparkOscope: Enabling Apache Spark Optimization through Cross Stack Monitoring - Yiannis Gkoufas
-
29:43
A Predictive Analytics Workflow on DICOM Images using Apache Spark
-
29:47
Natural Language Processing with CNTK and Apache Spark - Ali Zaidi
-
34:59
NLP with MLlib: Global Empire Building for Fun and Profit - Michelle Casbon