Designing a Horizontally Scalable Event Driven Big Data Architecture w/ Apache Spark Ricardo Fanjul
2,596 views · Published 11 October 2018 · 30:51 · Indexed 2 October 2026
Channel: Databricks · 2018 · Science & Technology
Traditional data architectures are not enough to handle the huge amounts of data generated from millions of users. In addition, the diversity of data sources are increasing every day: Distributed file systems, relational, columnar-oriented, document-oriented or graph Databases. Letgo has been growing quickly during the last years. Because of this, we needed to improve the scalability or our data platform and endow it further capabilities, like “dynamic infrastructure elasticity”, real-time processing or real-time complex event processing. In this talk, we are going to dive deeper into our journey. We started from a traditional data architecture with ETL and Redshift, till nowadays where we successfully have made an event oriented and horizontally scalable data architecture. We will explain in detail from the event ingestion with Kafka / Kafka Connect to its processing in streaming and batch with Spark. On top of that, we will discuss how we have used Spark Thrift Server / Hive Metastore as glue to exploit all our data sources: HDFS, S3, Cassandra, Redshift, MariaDB … in a unified way from any point of our ecosystem, using technologies like: Jupyter, Zeppelin, Superset ⦠We will also describe how to made ETL only with pure Spark SQL using Airflow for orchestration. Along the way, we will highlight the challenges that we found and how we solved them. We will share a lot of useful tips for the ones that also want to start this journey in their own companies. About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business. Read more here: https://databricks.com/product/unified-data-analytics-platform Connect with us: Website: https://databricks.com Facebook: https://www.facebook.com/databricksinc Twitter: https://twitter.com/databricks LinkedIn: https://www.linkedin.com/company/databricks Instagram: https://www.instagram.com/databricksinc/ Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-named-leader-by-gartner
More from this channel
-
10:06
Big Data Meets Learning Science
-
30:18
Yelp Ad Targeting at Scale with Apache Spark - Inaz Alaei-Novin and Joe Malicki
-
37:42
Scaling Data Science Capabilities with Apache Spark at Stitch Fix - Derek Bennett
-
30:50
Lazy Join Optimizations Without Upfront Statistics - Matteo Interlandi
-
31:04
Debugging Big Data Analytics in Apache Spark with BigDebug Matteo Interlandi and Muhammad Ali Gulzar
-
28:55
Scaling Up Data Science Applications with Kexin Xie and Yacov Salomon
-
31:03
Extending Spark Machine Learning: Adding Your Own Algorithms and Tools
-
23:09
The Key to Machine Learning is Prepping the Right Data - Jean Georges Perrin