Sergii Khomenko - From Data Science to Production - deploy, scale, enjoy!
2,948 views · Published 26 March 2016 · 39:28 · Indexed 22 September 2026
Channel: PyData · 2016 · Science & Technology
PyData Amsterdam 2016 Description Data cleaning is the first step of every Data Science project. Next one does Data Science. The talk covers a missing step of deployment and scaling Data Applications in production. We will go through all major steps of the process like Dockerizing application, Continuous Deployment with further AWS stack creation and rolling deploys although also covering new trends in Serverless architecture. Abstract Data Science is quite a young field. One of the definitions of Data Scientist: Person who is better at statistics than any software engineer and better at software engineering than any statistician. Hence, it's quite important to talk not only about best practices of feature generation and not overfitting but also about more of software engineering topics. The talk is based on our experience of Data Science developments at Stylight, an international fashion e-commerce company, that operates in 15 countries worldwide. We refer to our Data Applications written in R and Python, Scala; but the content is not limited to mentioned languages and applicable others. The talk consists three main parts. A first part introduces best practices of development. How to structure your development, make deployment easy and reproducible, how to make Continuous Integration and commit triggered deployments. The second part covers production deployment to AWS stack, in particular focusing on concepts of immutable infrastructure and infrastructure as code. The last part about using serverless architecture for data applications. We introduce an example of our outlier detection system, that automatically scales based on such approach. 00:00 Welcome! 00:10 Help us add time stamps or captions to this video! See the description for details. Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: https://github.com/numfocus/YouTubeVideoTimestamps
More from this channel
-
41:50
Shawn Scully: Creating an intelligent world at Dato
-
24:46
Dino Viehland & Raymond Laghaeian: Jupyter Notebooks and ML Model Operationalization
-
34:32
Scott Sanderson: Developing an Expression Language for Quantitative Financial Modeling
-
44:04
Cornelia Levy-Bencheton (Keynote): The Pain of Discipline and the Power of Not Yet
-
38:29
Simon Byrne - Julia for data analysis
-
49:10
Rui Miguel Forte - The CV: A Data Scientist's View
-
28:01
Ruby Childs & Nick Sorros - Making Recommendations without Data
-
50:23
PyData London 2016 Lightning Talks and Closing Address