Efficiently Triaging CI Pipelines with Apache Spark (Ivan Jibaja)
722 views · Published 27 September 2018 · 30:28 · Indexed 21 September 2026
Channel: Databricks · 2018 · Science & Technology
Ivan Jibaja, a tech lead in Pure Storage, discusses the continuous integration (CI) pipelines generate massive amounts of messy log data. At Pure Storage engineering, we run over 65,000 tests per day creating a large triage problem. Spark’s flexible computing platform allows us to write a single application for both streaming and batch jobs to understand the state of our CI pipeline. Spark indexes log data for real-time reporting (Streaming), uses Machine Learning for performance modeling and prediction (Batch job), and re-indexes old data for newly encoded patters (Batch job). Learn more here: https://databricks.com/session/efficiently-triaging-ci-pipelines-with-apache-spark-mixing-52-billion-events-day-of-streaming-with-40-tb-hour-of-batch-processing Article you might like: https://databricks.com/session/apache-spark-based-hyper-parameter-selection-and-adaptive-model-tuning-for-deep-neural-networks About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business. Read more here: https://databricks.com/product/unified-data-analytics-platform Connect with us: Website: https://databricks.com Facebook: https://www.facebook.com/databricksinc Twitter: https://twitter.com/databricks LinkedIn: https://www.linkedin.com/company/databricks Instagram: https://www.instagram.com/databricksinc/ Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-named-leader-by-gartner
More from this channel
-
28:55
Scaling Up Data Science Applications with Kexin Xie and Yacov Salomon
-
31:03
Extending Spark Machine Learning: Adding Your Own Algorithms and Tools
-
16:23
SparkOscope: Enabling Apache Spark Optimization through Cross Stack Monitoring - Yiannis Gkoufas
-
19:53
Women in Big Data Lunch
-
33:31
Improving Traffic Prediction Using Weather Data - Ramya Raghavendra
-
30:15
Cardinality Estimation in Apache Spark 2.3 (Ron Hu & Zhenhua Wang)
-
31:45
Bringing an AI Ecosystem to the Domain Expert and Enterprise AI Developer (Frederick Reiss)
-
29:47
Time Series Anomaly Detection in Plaintext Using Apache Spark with Jerry Schirmer (SparkCognition)