Running Apache Spark on a High-Performance Cluster Using RDMA and NVMe Flash - Patrick Stuedi
1,840 views · Published 8 June 2017 · 30:35 · Indexed 3 October 2026
Channel: Databricks · 2017 · Science & Technology
Effectively leveraging fast networking and storage hardware (e.g., RDMA, NVMe, etc.) in Apache Spark remains challenging. Current ways to integrate the hardware at the operating system level fall short, as the hardware performance advantages are shadowed by higher layer software overheads. This session will show how to integrate RDMA and NVMe hardware in Spark in a way that allows applications to bypass both the operating system and the Java virtual machine during I/O operations. With such an approach, the hardware performance advantages become visible at the application level, and eventually translate into workload runtime improvements. Stuedi will demonstrate how to run various Spark workloads (e.g, SQL, Graph, etc.) effectively on 100Gbit/s networks and NVMe flash. About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business. Read more here: https://databricks.com/product/unified-data-analytics-platform Connect with us: Website: https://databricks.com Facebook: https://www.facebook.com/databricksinc Twitter: https://twitter.com/databricks LinkedIn: https://www.linkedin.com/company/databricks Instagram: https://www.instagram.com/databricksinc/ Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-named-leader-by-gartner
More from this channel
-
10:06
Big Data Meets Learning Science
-
30:18
Yelp Ad Targeting at Scale with Apache Spark - Inaz Alaei-Novin and Joe Malicki
-
37:42
Scaling Data Science Capabilities with Apache Spark at Stitch Fix - Derek Bennett
-
30:50
Lazy Join Optimizations Without Upfront Statistics - Matteo Interlandi
-
31:04
Debugging Big Data Analytics in Apache Spark with BigDebug Matteo Interlandi and Muhammad Ali Gulzar
-
24:50
Neuro Symbolic AI for Sentiment Analysis - Michael Malak
-
28:55
Scaling Up Data Science Applications with Kexin Xie and Yacov Salomon
-
31:03
Extending Spark Machine Learning: Adding Your Own Algorithms and Tools