Scaling Data Science Capabilities with Apache Spark at Stitch Fix - Derek Bennett
574 views · Published 8 June 2017 · 37:42 · Indexed 29 September 2026
Channel: Databricks · 2017 · Science & Technology
"At Stitch Fix, data scientists work on a variety of applications, including style recommendation systems, natural language processing, demand modeling and forecasting, and inventory analysis and recommendations. They've used Apache Spark as a crucial part of the infrastructure to support these diverse capabilities, running a large number of varying-sized jobs, typically 500-1,000 separate jobs each day - and often more. Their scaling problems are around capabilities and handling many simultaneous jobs, rather than raw data size. They have a large team of around 80 data scientists who own their pipelines from start to finish; there is no ""ETL engineering team"" that takes over. As a result, Stitch Fix has developed a self-service approach with their infrastructure to make it easy for the team to submit and track jobs, and they've grown an internal community of Spark users to help each other get started with the use of Spark. In this session, you'll learn about Stitch Fix's approach to using and managing Spark, including their infrastructure, execution service and other supporting tools. You'll also hear about the types of capabilities they support with Spark, how the team transitioned to using Spark, and lessons learned along the way. Stitch Fix's infrastructure utilizes many services from Amazon AWS, tools from Netflix OSS, as well as several home-grown applications. Session hashtag: #SFent4" About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business. Read more here: https://databricks.com/product/unified-data-analytics-platform Connect with us: Website: https://databricks.com Facebook: https://www.facebook.com/databricksinc Twitter: https://twitter.com/databricks LinkedIn: https://www.linkedin.com/company/databricks Instagram: https://www.instagram.com/databricksinc/ Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-named-leader-by-gartner
More from this channel
-
10:06
Big Data Meets Learning Science
-
30:18
Yelp Ad Targeting at Scale with Apache Spark - Inaz Alaei-Novin and Joe Malicki
-
30:50
Lazy Join Optimizations Without Upfront Statistics - Matteo Interlandi
-
31:04
Debugging Big Data Analytics in Apache Spark with BigDebug Matteo Interlandi and Muhammad Ali Gulzar
-
28:55
Scaling Up Data Science Applications with Kexin Xie and Yacov Salomon
-
31:03
Extending Spark Machine Learning: Adding Your Own Algorithms and Tools
-
23:09
The Key to Machine Learning is Prepping the Right Data - Jean Georges Perrin
-
30:16
Best Practices for Using Alluxio with Apache Spark - Cheng Chang & Haoyuan Li