FR
live
AI

IBM and NASA release an open-source foundation model for mapping the Moon

IBM and NASA open-source the NASA-IBM Lunar Foundation Model, a foundation model trained on decades of multi-instrument lunar data, paired with a unified 30-layer dataset. It cuts ice-deposit identification error by up to 22% and joins the Prithvi family.

A dark cratered lunar surface seen up close, a thin amber scan line tracing across the grey regolith.

September 10, 2026. IBM and NASA announce from Yorktown Heights the open-source release of the NASA-IBM Lunar Foundation Model, one of the first publicly available foundation models for scientific exploration of the Moon. Petabytes. Decades of sensor observations sit scattered and heterogeneous. 30 layers. The model rests on a unified dataset that spatially aligns more than thirty layers from nine instruments across four missions. Why it matters: foundation AI is stepping out of chat and code into planetary scientific discovery — open source, on Hugging Face.

A foundation model that does not write text

In 2026 the phrase “foundation model” conjures large language models. The NASA-IBM Lunar Foundation Model is their opposite in use: it does not write, it reads the Moon’s surface. Trained on a lunar corpus built by IBM and NASA researchers, its job is what scientists did by hand — surfacing hidden relationships between data types and resolutions that no single instrument provides on its own.

The starting problem is concrete. To study the lunar surface, a researcher either sifts through maps and images manually or turns to task-specific machine-learning models that are often low resolution, computationally heavy, and not accurate enough for geographic analysis. A foundation model changes the method: instead of building a new algorithmic system for every scientific question, you start from a shared model and adapt it to the task. That is the logic of the Prithvi family, which the lunar model now joins after the geospatial, weather, and heliophysics models.

The dataset, a first of its kind

Half of the announcement is about the dataset, not the model. No public, unified set existed that brought multimodal, multi-resolution lunar data into a common framework suitable for modern machine learning. So IBM and NASA built the first of its kind: an “ML-ready” dataset aggregating more than 30 spatially-aligned layers from 9 instruments across 4 missions — NASA’s Lunar Reconnaissance Orbiter (LRO) and GRAIL, plus complementary data from the JAXA SELENE/Kaguya mission. Tens of thousands of images and maps describe geophysical properties of the surface and subsurface.

That integration work is as valuable as the model itself: it hands the lunar science community a multimodal view no one had the means to reconstruct alone, and a foundation others can build on.

Three tasks, three measured gains

The technical paper authored by IBM and NASA documents three use cases, each with a measurable gain over SwinV2-B, a reference model trained on ImageNet.

Ice deposits. Permanently shadowed regions are among the hardest to observe, yet they may hold ice below the surface — water and oxygen considered essential for a future Moon base and for producing rocket fuel toward Mars. The model reduces error (RMSE) in identifying high-potential ice areas by up to 22% versus SwinV2-B.

Volcanic history. The Irregular Mare Patches, lunar volcanic features, inform the Moon’s thermal evolution. Using imperfect labels, the model captures the extent of these features 3% better than SwinV2-B, at comparable accuracy but with lower fine-tuning cost.

Crater detection. Mapping craters helps NASA select safe landing sites, avoid steep slopes, and plan long-term infrastructure. At context scale (~100 m), the model outperforms SwinV2-B by nearly 19% using just half the training data.

The model is available on Hugging Face under the nasa-ibm-ai4science organization, alongside the technical report.

What open source changes for science

The open-source choice is not cosmetic. Kevin Murphy, NASA’s chief science data officer, puts it plainly: NASA spent decades building an extraordinary scientific record of the Moon, but collecting data is only part of the job — the rest is making it easy for scientists to explore and use. Juan Bernabe-Moreno, director of IBM Research Europe, frames the model as an open platform the global research community can extend, connecting observations across instruments to reveal patterns that are hard to see in isolation.

For research, the stakes go beyond the Moon. The Prithvi family — geospatial, weather, heliophysics, now the Moon — embodies a vision: stop building a new algorithm per scientific question, and pool a shared model that gets adapted. It is the economics of open-source software, transposed to observational science.

The technical trick: multimodal and multi-resolution

The model’s performance comes down to an architecture matched to the problem. Lunar observations are neither homogeneous nor at a common scale: a high-resolution LRO map, a GRAIL gravity measurement, and a Kaguya image describe different realities — surface, subsurface, composition. The NASA-IBM Lunar Foundation Model is built multimodal and multi-resolution: it aligns these sources into a common representation before reasoning over them, which explains the gains over a SwinV2-B model trained on terrestrial images.

The choice is not Moon-specific. It is the same principle behind IBM’s geospatial Prithvi model, used to monitor floods, crops, or deforestation on Earth. The Moon is a convenient testbed — abundant data, slowly changing terrain — but the method transfers to Earth-observation science, where the same heterogeneous-data problem exists at even greater scale.

What the model does not replace

The NASA-IBM Lunar Foundation Model is a foundation, not an oracle. It speeds up the reading of data, but scientific validation stays human: the training labels are imperfect, the announced gains are measured on specific tasks, and the model does not replace physics. Its role is to point a researcher toward what deserves study, not to decide.

For the community, the value lies elsewhere: by publishing the model with its dataset and technical report, IBM and NASA provide a reproducible base that any team can pick up, fine-tune, and critique — which is exactly the condition for science that moves forward. To try it, the entry point is the nasa-ibm-ai4science organization on Hugging Face, where the model, the technical report, and the unified dataset are published together.

The timing is not accidental. With NASA planning a sustained human return to the Moon, landing-site selection — where to set down, where ice is reachable, where the terrain is stable — is exactly the kind of question a foundation model can help answer faster. Opening the model now puts that tool in the hands of the broader planetary-science community before the mission hardware is fixed.

Verdict

The NASA-IBM Lunar Foundation Model is not a conversational model, and that is exactly the point: it shows a path where foundation AI creates value without being a chatbot, and where open source accelerates discovery rather than benchmark races.

If you work in remote sensing or geospatial, this model and its unified dataset are a reusable starting point, not a space curiosity. If you track the foundation-AI signal, note the direction: the next models that matter will observe, not converse. If you are skeptical of open scientific AI, this is a case where a model published with its dataset and paper moves an entire field faster than a closed one.

References

The cyber brief, every Tuesday

The flaws that matter and the patches to apply, in a ten-minute read.

No spam. One-click unsubscribe.
read next

On the same topic

Perplexity launches its local agent on Windows, gated behind 24 GB of VRAM

Perplexity has brought Portable Computer, the local edition of its Computer agent, to Windows after Linux and macOS — but only for NVIDIA RTX cards with at least 24 GB of VRAM. Simple tasks run on-device, the model hands off to the cloud when it needs more reasoning, and sensitive files can stay on the machine.

Four labs ship frontier models in one week and trigger model fatigue

In early September 2026, Anthropic, Meta, Google and OpenAI each ship a frontier model in the same week, and CNBC names the phenomenon model fatigue. A model’s real cost now depends on cache and context as much as benchmarks: stop comparing scores, compare the price of your workload.

← Back to the feed

Type at least two characters.

navigate open esc dismiss