Efficiently Triaging CI Pipelines with Apache Spark (Ivan Jibaja)

722 views · Published 27 September 2018 · 30:28 · Indexed 21 September 2026

Channel: Databricks · 2018 · Science & Technology

Watch on YouTube

Ivan Jibaja, a tech lead in Pure Storage, discusses the continuous integration (CI) pipelines generate massive amounts of messy log data. At Pure Storage engineering, we run over 65,000 tests per day creating a large triage problem. Spark’s flexible computing platform allows us to write a single application for both streaming and batch jobs to understand the state of our CI pipeline. Spark indexes log data for real-time reporting (Streaming), uses Machine Learning for performance modeling and prediction (Batch job), and re-indexes old data for newly encoded patters (Batch job).


Learn more here: https://databricks.com/session/efficiently-triaging-ci-pipelines-with-apache-spark-mixing-52-billion-events-day-of-streaming-with-40-tb-hour-of-batch-processing



Article you might like: https://databricks.com/session/apache-spark-based-hyper-parameter-selection-and-adaptive-model-tuning-for-deep-neural-networks
About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business.
Read more here: https://databricks.com/product/unified-data-analytics-platform

Connect with us:
Website: https://databricks.com
Facebook: https://www.facebook.com/databricksinc
Twitter: https://twitter.com/databricks
LinkedIn: https://www.linkedin.com/company/databricks
Instagram: https://www.instagram.com/databricksinc/ Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-named-leader-by-gartner

More from this channel