How Transfer Learning Bridges Data Gaps in Colorado Using California’s Rich Snow Data

Figure: Spatial extent and frequency of lidar-derived SWE maps. More snapshots of ASO in California support the transfer of learning to enhance SWE predictions in data-scarce Colorado. Axes show longitude and latitude.
The Science
This study employed a machine learning method called transfer learning to predict the amount of water stored in snow (known as Snow Water Equivalent or SWE) in Colorado. Since California has more detailed snow data, the scientists used that information to help make better predictions in Colorado, where data is limited. They found that both places share common features, such as elevation and snowfall. Using these shared features, they trained a computer model to learn from California’s data and apply it to Colorado. This improved the prediction accuracy in Colorado by 20% and made the results more reliable and less biased.
The Impact
This study demonstrates that machine learning can effectively address problems where data is scarce, particularly in mountainous regions where snow measurements are challenging to obtain. By employing smart techniques and utilizing knowledge of geography, the model can still make accurate predictions about snowfall. This helps scientists plan for water use, understand climate change, and predict natural hazards, such as floods. The method, called transfer learning, works well when guided by tools like factor analysis that help identify important patterns. This idea can also be applied to studying other aspects, such as weather, land, or water, in areas where data is limited.
Summary
This research addresses a significant gap: Colorado has limited snow data, whereas California has extensive data. Scientists hypothesized that snow patterns in both places are influenced by similar factors, such as elevation, temperature, and the amount of snow that has fallen. They used a method called exploratory factor analysis to confirm this idea.
Then, they utilized a machine learning model known as an Artificial Neural Network (ANN). They trained it using California’s rich data and then applied it to make predictions in Colorado—a method known as transfer learning. The best version of the model (called TL1W) worked by retaining general knowledge from California and updating only the parts specific to Colorado. They also gave more importance to key features like elevation and snowfall.
This made the model more accurate and improved its ability to handle limited data. It also avoided common errors by focusing on what matters most. Overall, the study demonstrates that combining smart computer models with scientific knowledge can aid in predicting environmental patterns, even when data is missing, which is beneficial for managing water resources, studying climate change, and protecting against natural hazards.
Contact
Dipankar Dwivedi
Lawrence Berkeley National Laboratory
Eoin L. Brodie, Watershed Function SFA LRM
Lawrence Berkeley National Laboratory
Funding
● OASIS (Open Source AI Software Infrastructure for Science); Supported by the U.S. Department of Energy, Office of Science, Office of Acquisition and Assistance
● ExaSheds Project; Supported by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research; Earth and Environmental Systems Sciences Division, Data Management Program
● Watershed Function Scientific Focus Area funded by the US Department of Energy, Office of Science, Office of Biological and Environmental Research
Publications
El Halabi, L., Mital, U., & Dwivedi, D. (2025). Modeling spatial distribution of snow water equivalent using transfer learning across mountainous basins. Journal of Geophysical Research: Machine Learning and Computation. https://doi.org/10.1029/2024JH000278
