Extending Spark Machine Learning: Adding Your Own Algorithms and Tools
4,095 views · Published 12 June 2017 · 31:03 · Indexed 21 September 2026
Channel: Databricks · 2017 · Science & Technology
Apache Spark's machine learning (ML) pipelines provide a lot of power, but sometimes the tools you need for your specific problem aren't available yet. This talk introduces Spark's ML pipelines, and then looks at how to extend them with your own custom algorithms. By integrating your own data preparation and machine learning tools into Spark's ML pipelines, you will be able to take advantage of useful meta-algorithms, like parameter searching and pipeline persistence (with a bit more work, of course). With Holden Karau and Seth Hendrickson About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business. Read more here: https://databricks.com/product/unified-data-analytics-platform Connect with us: Website: https://databricks.com Facebook: https://www.facebook.com/databricksinc Twitter: https://twitter.com/databricks LinkedIn: https://www.linkedin.com/company/databricks Instagram: https://www.instagram.com/databricksinc/ Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-named-leader-by-gartner
More from this channel
-
28:55
Scaling Up Data Science Applications with Kexin Xie and Yacov Salomon
-
16:23
SparkOscope: Enabling Apache Spark Optimization through Cross Stack Monitoring - Yiannis Gkoufas
-
19:53
Women in Big Data Lunch
-
33:31
Improving Traffic Prediction Using Weather Data - Ramya Raghavendra
-
30:28
Efficiently Triaging CI Pipelines with Apache Spark (Ivan Jibaja)
-
30:15
Cardinality Estimation in Apache Spark 2.3 (Ron Hu & Zhenhua Wang)
-
31:45
Bringing an AI Ecosystem to the Domain Expert and Enterprise AI Developer (Frederick Reiss)
-
29:47
Time Series Anomaly Detection in Plaintext Using Apache Spark with Jerry Schirmer (SparkCognition)