Productionizing H2O Models with Apache Spark (Kuba)
451 views · Published 20 September 2018 · 31:25 · Indexed 1 October 2026
Channel: Databricks · 2018 · Science & Technology
Jakub Hava (or "Kuba") is a cluster monitoring expert at H2O. Kuba shares about H2O's current Sparkling Water Project. Spark pipelines represent a powerful concept to support productionizing machine learning workflows. Their API allows to combine data processing with machine learning algorithms and opens opportunities for integration with various machine learning libraries. However, to benefit from the power of pipelines, their users need to have a freedom to choose and experiment with any machine learning algorithm or library. Therefore, we developed Sparkling Water that embeds H2O machine learning library of advanced algorithms into the Spark ecosystem and exposes them via pipeline API. We’ll demonstrate creation of pipelines integrating H2O machine learning models and their deployments using Scala or Python. To learn more: https://databricks.com/blog/2015/12/02/databricks-and-h2o-make-it-rain-with-sparkling-water.html About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business. Read more here: https://databricks.com/product/unified-data-analytics-platform Connect with us: Website: https://databricks.com Facebook: https://www.facebook.com/databricksinc Twitter: https://twitter.com/databricks LinkedIn: https://www.linkedin.com/company/databricks Instagram: https://www.instagram.com/databricksinc/ Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-named-leader-by-gartner
More from this channel
-
10:06
Big Data Meets Learning Science
-
30:18
Yelp Ad Targeting at Scale with Apache Spark - Inaz Alaei-Novin and Joe Malicki
-
37:42
Scaling Data Science Capabilities with Apache Spark at Stitch Fix - Derek Bennett
-
30:50
Lazy Join Optimizations Without Upfront Statistics - Matteo Interlandi
-
31:04
Debugging Big Data Analytics in Apache Spark with BigDebug Matteo Interlandi and Muhammad Ali Gulzar
-
28:55
Scaling Up Data Science Applications with Kexin Xie and Yacov Salomon
-
31:03
Extending Spark Machine Learning: Adding Your Own Algorithms and Tools
-
23:09
The Key to Machine Learning is Prepping the Right Data - Jean Georges Perrin