NASA and IBM release open lunar AI model for reproducible Moon science


IBM Newsroom
news
IBM and NASA Release Open-Source AI Model to Support Lunar Exploration
“IBM and NASA announced the open-source release of the NASA-IBM Lunar Foundation Model and said it was trained on a lunar observation dataset curated by IBM and NASA researchers.”
IBM Research
other
A rough guide for going back to the Moon
“IBM Research said the model integrates observations across modalities, viewing angles and spatial scales, and can be adapted to tasks such as crater mapping, ice prospecting and volcanic-feature analysis.”
Hugging Face
data
NASA-IBM Lunar Foundation Model (NASA-IBM LFM)
“The model card lists an Apache-2.0 license, a ViT-B encoder-decoder architecture, SomBench training data and limitations including that ice outputs are prospectivity estimates rather than measured ice.”
Moon model
NASA and IBM released an open-source lunar foundation model on September 10 for ice prospecting, crater detection and volcanic-feature mapping.
Large corpus
The SomBench pretraining corpus lists 963,609 WAC tiles and 1,000,113 NAC bundles across low- and high-resolution lunar tracks.
Open weights
The model card lists an Apache-2.0 license, ViT-B architecture details and documented benchmark procedures for downstream fine-tuning.
IBM and NASA on September 10 released the NASA-IBM Lunar Foundation Model, an open-source AI system designed to help scientists analyze decades of lunar observations across instruments, resolutions and missions as space agencies prepare for renewed activity on the Moon.1
The model supports three near-term science and exploration tasks: identifying areas with potential lunar ice, detecting and classifying craters, and mapping volcanic features known as irregular mare patches.1 IBM said it outperformed widely used methods by up to 23% on some lunar-feature identification tasks, though results vary by benchmark and include more modest gains for some applications.1
The release package may matter more than the headline performance. Alongside the model, NASA and IBM published SomBench, a machine-learning-ready lunar corpus built from spatially aligned observations that combine optical imagery, topography, slope, aspect and other scientific layers into common tile formats.4 For lunar science teams, that could reduce a persistent barrier: before testing a model, researchers often must reconcile observations collected at different resolutions, by different instruments and for different scientific purposes.2
The NASA-IBM model was trained on SomBench, a multimodal lunar tile corpus with two main tracks. The low-resolution track is anchored to Lunar Reconnaissance Orbiter Camera Wide Angle Camera, or WAC, visible data at 100 meters per pixel. The high-resolution track is anchored to Narrow Angle Camera, or NAC, imagery at 1 meter per pixel.4
The dataset card lists 963,609 WAC low-resolution tiles and 1,000,113 NAC high-resolution bundles, with 512-by-512-pixel tiles in both tracks.4 WAC tiles cover 51.2 kilometers by 51.2 kilometers each, while NAC tiles cover 512 meters by 512 meters each.4 The full pretraining data are hosted in an open AWS bucket, with samples and metadata available through Hugging Face.4
IBM and NASA said the broader unified dataset aggregates more than 30 spatially aligned layers from nine instruments across four missions, including NASA’s Lunar Reconnaissance Orbiter and GRAIL missions and Japan’s SELENE/Kaguya mission.1 The model card also credits data products from LROC, LOLA, Diviner, Mini-RF, Kaguya/SELENE, GRAIL, Lunar Prospector and the U.S. Geological Survey.3
That training mix matters because lunar observations are not just images. Some layers describe reflectance, terrain, thermal conditions or geophysical properties. Others exist only in polar or otherwise limited regions.4 The model’s scientific value depends on whether it can learn useful cross-modal relationships without erasing uncertainty introduced by resampling, missing coverage or inconsistent measurement conditions.
IBM and NASA present the system as a reusable foundation model rather than a task-specific detector. The public model card lists a ViT-B encoder-decoder architecture with 768-dimensional representations, 12 layers and 12 attention heads, pretrained for 150,000 steps on 16 Nvidia H100 GPUs.3
Researchers can fine-tune the model through TerraTorch. IBM said the team used lightweight LoRA adapters that kept most base weights frozen during adaptation.2 The approach is meant to lower the compute needed to adapt the model to new lunar tasks, a practical issue for research groups that cannot train large vision models from scratch.
On benchmarks, NASA and IBM report the strongest headline result for potential ice mapping. IBM’s announcement said the model reduced root-mean-square error by up to 22% against a SwinV2-B ImageNet baseline when identifying areas with high potential for lunar ice.1 The IBM Research explainer describes the task as especially difficult because polar craters can remain in permanent shadow and because ice-relevant signals are distributed across temperature, terrain and other measurements.2
For crater detection, the model achieved accuracy comparable to SwinV2-B at meter-scale resolution while offering greater efficiency and lower fine-tuning costs, according to IBM.1 At about 100-meter context-scale resolution, IBM said it outperformed SwinV2-B by nearly 19% while using half the training data.1
For volcanic features, the gains are narrower. IBM said the model mapped irregular mare patches 3% better than a SwinV2-B baseline when trained with imperfect labels, while the IBM Research post described the result as comparable accuracy at lower cost.12
The release is unusually explicit about limitations, an important point in a domain where a visually plausible output can still be scientifically misleading. The model card says the ice-prospectivity output regresses a knowledge-driven fuzzy-overlay map rather than directly measured ice.3 In other words, the model is not detecting buried water ice as a physical instrument would. It is learning to reproduce or improve a proxy map built from existing assumptions and measurements.
The model card also notes that qualitative generation tests are sanity checks rather than systematic evaluations, and that absolute values can drift.3 It further says the contribution of geometry tokenization, mixed-resolution training and lunar pretraining has not yet been fully separated through ablation studies.3
Coverage is another constraint. The NAC pretraining set is limited to sites with co-registered 3-meter stereo digital terrain models, which the card describes as globally distributed but not globally dense.3 As a result, the model may be strongest where training data are richest and less reliable in areas with sparse coverage or different observing conditions.
To reduce leakage between training and evaluation, SomBench assigns splits by lunar grid cell. Train, validation and test zones are separated geographically, and boundary-straddling tiles are dropped.4 That is a meaningful reproducibility step because neighboring lunar tiles can be highly correlated. Random tile-level splits could make a model appear more general than it is.
The open release appears to lower several barriers at once: model weights are available under an Apache-2.0 license, the architecture and training setup are documented, benchmark loaders are provided, and the underlying pretraining corpus is exposed through public infrastructure.34 Those are not trivial details for scientific AI, where reproducibility depends on more than a downloadable checkpoint.
The Next Web reported that the dataset may be as important as the model because lunar observations have historically been stored in different formats and resolutions, complicating machine-learning work.7 That assessment aligns with the NASA-IBM framing: the release is not just a larger model, but a common framework for aligning lunar measurements that were previously hard to combine.1
Still, open access does not make the system mission-ready. Teams preparing for Artemis-era science will need to validate outputs against independent observations, quantify uncertainty and decide when a model prediction is suitable for triage rather than evidence. A predicted ice-prospectivity map, for example, could help prioritize regions for analysis, but it cannot replace ground truth from instruments or landers.
The release marks a shift from AI demonstrations toward operational scientific tooling. In this setting, scale is only part of the story. The key questions are whether the data lineage is clear, whether benchmarks prevent leakage, whether uncertainty is communicated, and whether outside teams can reproduce or challenge the results.
For now, NASA and IBM have supplied a stronger foundation than a closed model announcement would have: open code paths, documented data splits, public weights and a large aligned corpus. Whether it becomes useful lunar infrastructure will depend on how quickly independent planetary scientists can test it on questions the original team did not choose.

Apple’s first foldable iPhone is not just a hardware catch-up to Android rivals. Its larger test is whether iOS 27.1 can make a 7.6-inch folding screen feel like a flexible mobile workspace without turning the device into a small iPad.

ETH Zurich’s Swiss National Supercomputing Centre will host Switzerland’s first IBM Quantum System Two, giving Swiss researchers and selected companies a dedicated route into IBM’s Nighthawk-based hardware. The near-term value will depend less on headline qubit counts than on how users combine quantum circuits with classical supercomputing workflows.

CISA’s September 10 update links CVE-2025-14733 to known ransomware use, turning an old WatchGuard Firebox patch issue into an incident-response priority. Security teams should verify exposure, patch or retire vulnerable devices, and check internet-facing firewalls for signs of compromise.

CoreWeave’s new Physical AI Field Engineering service embeds domain specialists with enterprise engineering teams to turn proprietary test, simulation and telemetry data into production AI. The offering is credible where customers need help operationalizing models against real-world physics, but it also deepens CoreWeave’s role as a high-touch services layer around its AI cloud infrastructure.
Foundation model
A model pretrained on broad data so it can be adapted to multiple downstream tasks instead of being built for only one purpose.
SomBench
NASA and IBM’s lunar machine-learning dataset of co-registered image tiles and scientific layers used for pretraining and benchmarking.
LoRA
Low-rank adaptation, a fine-tuning method that updates small adapter modules while leaving most base-model weights unchanged.
Prospectivity map
A map estimating where a resource, such as lunar ice, may be likely; it is not the same as direct measurement of that resource.
Comments