Stateful Structure Streaming and Markov Chains Join Forces to Monitor the Biggest Storage of Physics

141 views · Published 10 October 2018 · 28:44 · Indexed 28 September 2026

Channel: Databricks · 2018 · Science & Technology

Watch on YouTube

Talk by Daniel Lanza (CERN)
Many organisations face the difficult challenge of enabling Machine Learning projects to get to market more quickly and to allow data science teams to share their feature. In this talk, I will be discussing the machine learning pipeline developed at one of Australia's largest telecommunications companies to achieve this goal using Spark and Spark ML as well as the challenges faced along the way. I'll begin by discussing the utility and motivation for a centralised feature store, before looking at the complexities of such an undertaking (both technical and otherwise). We will then dig into the technical details of implementation by discussing the scalability headaches we faced and dive into the details of the solutions used to drastically improve the speed and organisational scalability of the system. Several areas that will be covered are providing a declarative API that allowed us to compile feature definitions into optimised spark code, tuning the spark DAG for drastically improved parallelism, adjusting the workflow for different machine learning use cases and fine tuning the resource allocation to avoid unnecessary bottlenecks. Finally we will touch on lessons learnt along the way and offer advice on things to avoid as well as how to take things to the next level. Attendees will leave with a better understanding of the challenges faced with scaling a centralised machine learning pipeline at a large organisation. They will also have plenty of real world practical optimisation techniques that they can apply to their own problems.

About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business.
Read more here: https://databricks.com/product/unified-data-analytics-platform

Connect with us:
Website: https://databricks.com
Facebook: https://www.facebook.com/databricksinc
Twitter: https://twitter.com/databricks
LinkedIn: https://www.linkedin.com/company/databricks
Instagram: https://www.instagram.com/databricksinc/ Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-named-leader-by-gartner

More from this channel