Scaling Genomics Pipelines in the Cloud (Ram Sriharsha & Frank Austin Nothaft)
135 views · Published 27 September 2018 · 30:01 · Indexed 23 September 2026
Channel: Databricks · 2018 · Science & Technology
Ram Sriharsha, a Product Manager at Databricks, and Frank Austin Nothaft, a masters of Science in Computer Science from UC Berkeley, discuss how next generation sequencing is becoming cheaper and more accessible. The volume of data sequenced is increasing faster than Moore’s Law. However, it is still expensive and slow to go from raw reads to variant calls, and to produce annotated variants that can then be analyzed downstream. In this talk, we will discuss the first state of the art, scalable and simple DNA sequencing workflow that is built on top of Apache Spark and the Databricks APIs. The pipeline is simple to set up, is easy to scale out, and can sequence a 30x coverage genome cost efficiently on the cloud. Learn more here: https://databricks.com/session/serverless-machine-learning-on-modern-hardware-using-apache-spark Article you might like: https://databricks.com/session/dlobd-an-emerging-paradigm-of-deep-learning-over-big-data-stacks About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business. Read more here: https://databricks.com/product/unified-data-analytics-platform Connect with us: Website: https://databricks.com Facebook: https://www.facebook.com/databricksinc Twitter: https://twitter.com/databricks LinkedIn: https://www.linkedin.com/company/databricks Instagram: https://www.instagram.com/databricksinc/ Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-named-leader-by-gartner
More from this channel
-
10:06
Big Data Meets Learning Science
-
28:55
Scaling Up Data Science Applications with Kexin Xie and Yacov Salomon
-
31:03
Extending Spark Machine Learning: Adding Your Own Algorithms and Tools
-
16:23
SparkOscope: Enabling Apache Spark Optimization through Cross Stack Monitoring - Yiannis Gkoufas
-
34:59
NLP with MLlib: Global Empire Building for Fun and Profit - Michelle Casbon
-
19:53
Women in Big Data Lunch
-
29:48
Deep Dive into Deep Learning Pipelines continues - Sue Ann Hong & Tim Hunter
-
33:31
Improving Traffic Prediction Using Weather Data - Ramya Raghavendra