Stateful Structure Streaming and Markov Chains Join Forces to Monitor the Biggest Storage of Physics
141 views · Published 10 October 2018 · 28:44 · Indexed 28 September 2026
Channel: Databricks · 2018 · Science & Technology
Talk by Daniel Lanza (CERN) Many organisations face the difficult challenge of enabling Machine Learning projects to get to market more quickly and to allow data science teams to share their feature. In this talk, I will be discussing the machine learning pipeline developed at one of Australia's largest telecommunications companies to achieve this goal using Spark and Spark ML as well as the challenges faced along the way. I'll begin by discussing the utility and motivation for a centralised feature store, before looking at the complexities of such an undertaking (both technical and otherwise). We will then dig into the technical details of implementation by discussing the scalability headaches we faced and dive into the details of the solutions used to drastically improve the speed and organisational scalability of the system. Several areas that will be covered are providing a declarative API that allowed us to compile feature definitions into optimised spark code, tuning the spark DAG for drastically improved parallelism, adjusting the workflow for different machine learning use cases and fine tuning the resource allocation to avoid unnecessary bottlenecks. Finally we will touch on lessons learnt along the way and offer advice on things to avoid as well as how to take things to the next level. Attendees will leave with a better understanding of the challenges faced with scaling a centralised machine learning pipeline at a large organisation. They will also have plenty of real world practical optimisation techniques that they can apply to their own problems. About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business. Read more here: https://databricks.com/product/unified-data-analytics-platform Connect with us: Website: https://databricks.com Facebook: https://www.facebook.com/databricksinc Twitter: https://twitter.com/databricks LinkedIn: https://www.linkedin.com/company/databricks Instagram: https://www.instagram.com/databricksinc/ Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-named-leader-by-gartner
More from this channel
-
10:06
Big Data Meets Learning Science
-
30:18
Yelp Ad Targeting at Scale with Apache Spark - Inaz Alaei-Novin and Joe Malicki
-
28:55
Scaling Up Data Science Applications with Kexin Xie and Yacov Salomon
-
31:03
Extending Spark Machine Learning: Adding Your Own Algorithms and Tools
-
23:09
The Key to Machine Learning is Prepping the Right Data - Jean Georges Perrin
-
30:16
Best Practices for Using Alluxio with Apache Spark - Cheng Chang & Haoyuan Li
-
27:41
Social Media, Spark, Machine Learning, and Data Visualization to Find Patterns and Insight
-
16:23
SparkOscope: Enabling Apache Spark Optimization through Cross Stack Monitoring - Yiannis Gkoufas