How I Learned to Stop Worrying and Love the Exascale

1,778 views · Published 6 September 2018 · 1:12:16 · Indexed 29 September 2026

Channel: Microsoft Research · 2018 · Science & Technology

Watch on YouTube

Our ability to produce data continues to grow exponentially, and so does the computational power available in the world. By 2021, it is expected that zettabytes will be stored in the cloud, and the DoE will have deployed heterogeneous exascale clusters with millions of CPU cores and tens of thousands of nodes. Resources at this scale have the potential to accelerate scientific breakthroughs, but for that to happen we need to build scalable software systems that will utilize them efficiently.

Today, these clusters confine applications to operate in a non-privileged environment without feedback. I will be discussing our two-pronged effort to improve operational efficiency in this setting by: (a) allowing applications to control mechanisms and observe information traditionally accessible only to the operating system, and (b) using machine learning models to inform users of expected job behavior. This work is a collaboration with the Los Alamos and Argonne National Labs.

See more at https://www.microsoft.com/en-us/research/video/how-i-learned-to-stop-worrying-and-love-the-exascale/

More from this channel