Uwe L Korn - Efficient and portable DataFrame storage with Apache Parquet
2,931 views · Published 15 May 2017 · 28:31 · Indexed 21 September 2026
Channel: PyData · 2017 · Science & Technology
Filmed at PyData London 2017 www.pydata.org Description Apache Parquet is the most used columnar data format in the big data processing space and recently gained Pandas support. It leverages various techniques to store data in a CPU and I/O efficient way and provides capabilities to push-down queries to the I/O layer. In this talk, it is shown how to use it in Python, detail its structure and present the portable usage with other tools. Abstract Since its creation in 2013, Apache Parquet has risen to be the most widely used binary columnar storage format in the big data processing space. While supporting basic attributes of a columnar format like reading a subset of columns, it also leverages techniques to store the data efficiently while providing fast access. In addition the format is structured in such a fashion that when supplied to a query engine, Parquet provides indexing hints and statistics to quickly skip over chunks of irrelevant data. In recent months, efficient implementations to load and store Parquet files in Python became available, bringing the efficiency of the format to Pandas DataFrames. While this provides a new option to store DataFrames, it especially allows us to share data between Pandas and a lot of other popular systems like Apache Spark or Apache Impala. In this talk we will show the improvements that Parquet bring performance-wise but also will highlight important aspects of the format that make it portable and efficient for queries on large amount of data. As not all features are yet available in Python, an overview of the upcoming Python-specific improvements and how the Parquet format will be extended in general is given at the end of the talk. PyData is an educational program of NumFOCUS, a 501(c)3 non-profit organization in the United States. PyData provides a forum for the international community of users and developers of data analysis tools to share ideas and learn from each other. The global PyData network promotes discussion of best practices, new approaches, and emerging technologies for data management, processing, analytics, and visualization. PyData communities approach data science using many languages, including (but not limited to) Python, Julia, and R. We aim to be an accessible, community-driven conference, with novice to advanced level presentations. PyData tutorials and talks bring attendees the latest project features along with cutting-edge use cases. 00:00 Welcome! 00:10 Help us add time stamps or captions to this video! See the description for details. Want to help add timestamps to our YouTube videos to help with discoverability? Find out more here: https://github.com/numfocus/YouTubeVideoTimestamps
More from this channel
-
24:46
Dino Viehland & Raymond Laghaeian: Jupyter Notebooks and ML Model Operationalization
-
34:32
Scott Sanderson: Developing an Expression Language for Quantitative Financial Modeling
-
38:29
Simon Byrne - Julia for data analysis
-
49:10
Rui Miguel Forte - The CV: A Data Scientist's View
-
50:23
PyData London 2016 Lightning Talks and Closing Address
-
35:34
Frank Kaufer - Building a polyglot Data Science Platform on Big Data systems
-
33:19
Josh Yudaken | Building domain specific databases using Python for prototypes during take off
-
38:04
Peter Wang | Keynote: Python for Pythonistas