LEAP Momentum Bootcamp in Climate Data Science

Develop research skills, collaborate with experts, and accelerate your career through immersive Earth system science training.

Closed for Enrollment (Subscribe for Updates)

Modules/Weeks

1

Weekly Effort

2 days

Format

Cost

$100.00

Course Description

Enrollment for this course is currently closed. Join the waitlist below to receive an alert as soon as the next session opens for enrollment.


In the interdisciplinary LEAP Momentum Bootcamp in Climate Data Science, participants from diverse backgrounds will learn about machine learning, Python, deep learning, and climate modeling through hands-on, collaborative activities and mentorship, designed to accelerate their academic and career trajectories.

Key takeaways include:

  • Developing core research skills in climate and Earth system science through guided training and exercises.
  • Gaining practical experience with computational tools, data analysis, and scientific communication.
  • Building connections with faculty mentors and peers across diverse institutions and disciplines.
  • Strengthening career readiness through workshops on proposal writing, publishing, and science communication.


Who it’s for:

This live online bootcamp is designed for undergraduate and graduate students, as well as early-career researchers pursuing Earth system science. It is ideal for learners seeking to strengthen research skills, expand professional networks, and accelerate academic or career development.

This climate data science bootcamp welcomes:

  • Faculty members and research scientists from LEAP institutions (Columbia, NYU, University of Minnesota, University of California at Irvine, NASA/GISS, NCAR)
  • Postdocs and PhD students
  • Research scientists
  • NYC Public School teachers and Summer Institute alumni
  • LEAP partners
  • Members of the general public

Course Prerequisites

To make the most out of this experience, participants are encouraged to familiarize themselves with foundational concepts in machine learning, climate science, and introductory Python programming. We encourage you to review our learning resources on climate science and data science before the bootcamp. You do not need to study all the materials thoroughly—browse through them to become familiar with key vocabulary and concepts.

Participants should have:

  • Basic to intermediate experience in Python
  • A computer with reliable internet access
  • Comfort using standard digital collaboration tools (e.g., Zoom, shared documents, and basic data analysis platforms)

No specialized software or prior technical setup is required beyond what will be provided during the program.

What You Will Learn

By the end of this bootcamp, learners will be able to:

  • Apply core research methods in Earth system science, including data analysis, modeling, and interpretation.
  • Communicate scientific findings effectively through writing, presentations, and visualizations tailored to diverse audiences.
  • Collaborate across disciplines, integrating perspectives and approaches to address complex environmental questions.
  • Navigate academic and professional pathways with enhanced skills in proposal development, publishing, and career planning.
  • Discover, access, and explore open-access climate datasets, including satellite observations and climate simulations, using the Xarray Python package.
  • Calculate common climate statistics and diagnostics of variability and change using Xarray.
  • Perform interactive visualization of climate data using the Holoviews package.
  • Perform machine learning on spatio-temporal climate data.
  • Compare the performance and predictive skills of machine learning models.
  • Perform open science in the cloud using the LEAP-Pangeo Jupyter Hub.
     

Session Overviews

Session 1: Multi-dimensional Climate Data — From Excel to Xarray | Monday, January 12th: 9:00 - 12:00 EST

Learn how climate data differ from standard tabular formats and explore their multi-dimensional structure. This session introduces the powerful Python package Xarray, demonstrating its functionality through familiar Excel concepts. Participants will practice navigating and manipulating climate datasets efficiently using Xarray’s core tools.

Session 2: Climate Data Visualization and Processing Using Xarray | Monday, January 12th: 13:00-15:00 EST

Dive deeper into how climate data are organized, visualized, and processed in Python. This session continues building on Xarray fundamentals, showing how it integrates seamlessly with other core packages such as NumPy and Pandas to support robust and scalable climate data workflows.

Session 3: Introduction to Climate Models and the CMIP Dataset | Monday, January 12th: 15:20-17:00 EST

Gain an overview of climate models—what they are, how they work, and how scientists use them to understand and project Earth’s climate. This session introduces the Coupled Model Intercomparison Project (CMIP) dataset, which informs IPCC reports, and demonstrates how to access and analyze CMIP outputs using the LEAP-Pangeo platform.

Session 4: Introduction to Machine Learning for Climate Data Science | Tuesday, January 13th: 09:00–12:00 EST

Explore how machine learning complements traditional climate science. Through practical examples, you’ll learn key concepts such as supervised vs. unsupervised learning, model training and validation, and the challenges posed by high-dimensional, spatio-temporal climate data. Participants will practice preparing datasets and interpreting simple ML results in physically meaningful ways.

Session 5: Deterministic Machine Learning Methods | Tuesday, January 13th: 13:00–15:00 EST

This session focuses on deterministic machine learning approaches that learn explicit mappings between inputs and outputs—such as regression, neural networks, and hybrid physics-informed models. Using examples like temperature forecasting and parameter calibration, participants will discuss how these models approximate physical processes, the importance of interpretability, and how to prevent overfitting.

Session 6: Generative Machine Learning Methods | Tuesday, January 13th: 15:30–17:00 EST

Discover probabilistic and generative frameworks—including variational autoencoders (VAEs), normalizing flows, and diffusion models—that capture full distributions of climate variables rather than single deterministic outcomes. This session highlights how these methods support uncertainty quantification, emulate ensemble simulations, and reveal the evolving patterns of extreme-event distributions.

Instructors

Qingyuan Yang
Qingyuan Yang
Associate Research Scientist | Earth + Environmental Engineering, Columbia University

Qingyuan Yang is an Associate Research Scientist at Columbia University. His work focuses on applying machine learning methods to calibrate climate model parameters and improve model performance. He has also worked on volcanic ash fall hazard analysis. Before joining Columbia University, Yang was a postdoctoral researcher at Nanyang Technological University. His research interests include machine learning-based emulator development, parameter estimation, and uncertainty quantification for climate and geophysical models.

Candace Agonafir
Candace Agonafir
Associate Research Scientist | Data Science + Civil Engineering, Columbia University

Candace Agonafir has a Ph.D. in Civil Engineering from the City College of New York, a M.S. in Industrial Engineering from NYU, and a B.S. in Physical Science with a concentration in Physics and a minor in Catholic Theology from St. John’s University. Her past research involved using regression and machine learning methodologies to provide invaluable insights for advancing urban flooding detection, prediction, and prevention. Ultimately, by continuously adopting improved, contemporary machine learning techniques, she aims to become an established, impactful, and consistent contributor in the field of civil engineering water resources via Artificial Intelligence.

Shuolin Li
Shuolin Li
Postdoctoral Research Scientist | Data Science Institute, Columbia University

Shawn Li is a research scientist at the Data Science Institute and the LEAP at Columbia University. His research bridges fluid mechanics, hydrology, and machine learning to advance probabilistic data assimilation and climate risk modeling. He develops generative frameworks for predicting and interpreting non-Gaussian extremes—such as heatwaves and sediment transport events—by integrating physical theory with deep probabilistic models. Li received his Ph.D. in Fluid Dynamics and an M.S. in Computer Science from Duke University, and before that, earned graduate degrees from Cornell University and Northwestern University.

Course Dates

-
Subscribe for Course Updates
CAPTCHA