NASA and IBM Release Open Source Lunar Foundation Model
Science

NASA and IBM Release Open Source Lunar Foundation Model

TechNews Editorial
TechNews EditorialOct 4, 2026 · 2 min read
Share

Why it matters

The model makes decades of complex lunar observation data usable for machine learning, improving tasks like polar ice deposit prediction and crater detection for future research.

The facts

  • NASA and IBM released an open-source lunar foundation model using 17 years of orbiter data to aid machine learning.
  • The model reduced prediction error for polar ice deposits by up to 22 percent compared to the best baseline, according to IBM.
  • The model is publicly available on Hugging Face, with code hosted on GitHub and integrated into TerraTorch.

NASA and IBM Research, alongside several academic institutions, have released the NASA-IBM Lunar Foundation Model. The release creates an open source foundation model designed to make decades of lunar observation data usable for machine learning. Kevin Murphy, NASA's chief science data officer, noted that while NASA has spent decades building an extraordinary scientific record of the Moon, collecting data is only part of the job. The data must also become easier for scientists to use.

Unlike task-specific algorithms, foundation models undergo pretraining on large volumes of unlabeled data. They then adapt to specific tasks using only a few labeled examples. This design benefits lunar research, where observation data is plentiful but labels are scarce. The team trained the model from scratch using SomBench, which they described as the largest co-registered multimodal lunar corpus to date. The dataset contains nearly two million tile bundles across eleven modalities and two spatial scales.

Multimodal data powers model training

The collection includes about one million high-resolution images from the Narrow Angle Camera at roughly one meter per pixel. It also features nearly 964,000 multispectral images from the Wide Angle Camera at 100 meters per pixel. The bulk of the information comes from 17 years of observations by the Lunar Reconnaissance Orbiter. According to NASA, its data volume exceeds that of all other NASA planetary missions combined. Data from the GRAIL mission, Lunar Prospector, and JAXA's Kaguya probe completed the collection, bringing together over 30 spatially aligned data layers from nine instruments and four missions.

The architecture relies on TerraMind, a multimodal Earth observation model. Researchers trained the lunar model from scratch instead of fine-tuning an existing version. The system receives imaging geometry as explicit context for each tile, including illumination angles, sun position, and tile extent. The team also used FlexiViT to allow the trained model to adapt to tasks with different image patch sizes without retraining.

Read nextAnthropic's Mythos model uncovers Rejetto HTTP File Server flaw now under active attack

Model performance across specific tasks

The team tested the model on crater detection at 100-meter and 1-meter scales, polar ice deposit prediction, and segmenting Irregular Mare Patches. According to the technical report, the pretrained model matched or beat common baselines and an architecturally identical control model with random initialization. The largest gains occurred in ice deposit prediction. The model cut prediction error by up to 22 percent compared to the best baseline, SwinV2-B, according to IBM. For coarse-scale crater detection, the model beat SwinV2-B by nearly 19 percent while using only half the data.

The research shows the model is not suited for absolute geodetic positioning. In generation tests, latitude and longitude were off by dozens of degrees in some cases. Elevation structures also appeared with shifted absolute height values despite correct shape reconstruction. The authors view the release as a reusable foundation for downstream tasks rather than a replacement for physical measurement instruments.

The model is now publicly available on Hugging Face. The code is hosted on GitHub and integrated into the open-source toolkit TerraTorch. The team also released the machine learning ready pretraining datasets and benchmark collections. Both organizations developed the model as part of their ongoing Space Act Agreement collaboration.

Newsletter

Get the best AI & tech news daily

A concise daily digest. Unsubscribe anytime.

We use your email only to send this newsletter.

Keep reading