Connect with us

NEWS

NASA and IBM’s Lunar Model Cuts Polar Ice-Map Error

NASA and IBM’s open lunar foundation model cuts polar ice-map error by up to 22%, while its own card says it does not measure ice or certify landing sites.

Published

on

NASA and IBM released an open lunar foundation model on September 10, 2026, that cut polar ice-map error by up to 22%. The NASA-IBM Lunar Foundation Model was trained on nearly 2 million co-registered tile bundles, mostly from 17 years of Lunar Reconnaissance Orbiter observations, and posted to Hugging Face under Apache-2.0.

The public pitch is a reusable Moon map for craters, young volcanic patches, and polar ice. The artifact is narrower: its strongest gain is a match to an ice-likelihood overlay, and the card that ships with the weights says it is not a stand-in for instruments, coordinates, or landing-site calls.

What the Ice Score Measures

IBM and NASA evaluated the model on polar ice prospectivity regression, crater detection at 100-metre and 1-metre scales, and segmentation of irregular mare patches. The ice task is where the pretrained checkpoint pulled away from ImageNet baselines and from a random-init copy of the same architecture.

Full fine-tuning reached a root-mean-square error of 0.0293 on that ice map, against 0.0377 for SwinV2-B. IBM described that gap as a cut of up to 22% versus the SwinV2-B (ImageNet) baseline. A random-initialisation control scored 0.0397 and still beat five of six ImageNet baselines, which is why the paper says part of the ice gain is the way the model treats each data layer as its own token stream, not pretraining alone.

Kevin Murphy, NASA’s chief science data officer and acting chief data and AI officer, put the release in NASA’s lunar foundation model announcement as a data-access problem, not a new orbiter.

NASA has spent decades building an extraordinary scientific record of the Moon, but collecting data is only part of the job. We also have to make data easier for scientists to explore and use. The NASA-IBM Lunar Foundation Model shows what’s possible when we bring AI to NASA’s petabytes of scientific data. That’s a real opportunity we see with AI: turning large-scale data into new discoveries.

Kevin Murphy, chief science data officer, NASA Headquarters

That ice number is easy to over-read. The Hugging Face card says ice-prospectivity outputs regress a knowledge-driven fuzzy-overlay map, not measured ice. The model combines temperature and terrain, including slope, aspect, ice-stability depth, and maximum surface temperature, then paints likelihood. NASA’s own figure for the task shows Mons Mouton near the south pole, with the new model keeping fine prospectivity patterns that a ConvNeXt run smoothed away.

Polar Ice Is the Constraint on a Lunar Stay

NASA has aimed a sustained lunar stay at the south pole because permanently shadowed craters stay cold enough to hold water ice, while nearby ridges can sit in sunlight that charges solar arrays. IBM’s technical write-up is blunt about why that ice is on the task list: it could supply crews with drinking water, oxygen, and fuel.

A NASA radar instrument flown on India’s Chandrayaan-1 in 2008 found ice in 40 small craters near the north pole and put the water at 600 million metric tons, enough to fill at least 240,000 Olympic-sized swimming pools. Later work, including NASA’s retired SOFIA observatory, found water molecules stuck to grains of lunar dust. The architecture of a long stay still hangs on a simpler question than any of those detections: whether a crew can walk from sunlit ground into a cold trap that actually holds ice you can process.

NASA technical papers on Artemis in-situ resource use have already treated that geometry as a site filter. Julie Kleinhenz at NASA Glenn and colleagues scored south-pole regions of interest for ice mining, with the two regions near Shackleton ranking as more workable and de Gerlache presenting more trouble against the same criteria. Early mine sketches in that work focused on small, few-kilometre permanently shadowed regions, because a processing plant still needs nearby sun.

WHAT POLAR ICE IS FOR

  • Life support: Water for drinking and hygiene, plus oxygen split out by electrolysis.
  • Propellant: Hydrogen and oxygen produced on the surface instead of hauled from Earth.
  • Site geometry: Cold traps close enough to walk, next to high ground that stays in the sun.
  • Shelter: Polar crater floors and rims that can hide a crew from radiation and temperature swings.

Juan Bernabé-Moreno, director of IBM Research Europe for Ireland and the UK, described the model as a way to get the lay of the land before anyone sets out, and said he hoped it would help the next generation of astronauts find their way around. That is the operational reading the announcement invites. The weights that actually shipped still have to be read against the card’s bans, which is the next fact that matters for anyone treating this as a landing aid.

SomBench Aligns Nine Instruments Across Four Missions

The training set is SomBench, which the team calls the largest co-registered multimodal lunar corpus to date. It holds 1,000,113 Narrow Angle Camera bundles at about 1 metre per pixel and 963,609 Wide Angle Camera bundles at about 100 metres per pixel, a split that adds to 1,963,722 tiles, the nearly 2 million figure in the release. NASA says LRO’s archive is larger than all other NASA planetary missions combined and that the orbiter captured an almost seamless high-resolution mosaic of the whole Moon.

The stack is not camera-only. SomBench pulls more than 30 spatially aligned layers from nine instruments across four missions: LRO, GRAIL, Lunar Prospector, and JAXA’s Kaguya/SELENE. GRAIL mapped gravity at about 20 kilometres per pixel. LRO’s Narrow Angle Camera images boulders and crater rims at 1 metre per pixel. Kaguya adds mineral maps. Lunar Prospector supplies hydrogen abundance at the poles. Diviner, Mini-RF, and LOLA fill in heat, radar, and topography. Wide-angle tiles cover 51.2 kilometres on a side; narrow-angle tiles cover 512 metres.

The authors pretrained from scratch on SomBench, adapting IBM and ESA’s TerraMind masked-token recipe rather than fine-tuning the Earth checkpoint. Tiles are split by Lunar Transverse Mercator zones so overlapping geography cannot leak across train, validation, and test. NAC pretraining is globally spread but not globally dense: it is limited to 1,095 frames that have co-registered 3-metre stereo elevation models. Static-map context for polar-only instruments covers about 6,900 polar tiles per track.

Lighting Is Fed In as Its Own Input

On the Moon, lighting often does more to an image than the surface does. A lunar day is two weeks of sun and two weeks of dark. With almost no atmosphere to scatter light, peaks glare and valleys drop into black. IBM puts noon heat near 250°F (121°C) and shadowed crater floors near -410°F (-246°C). At the poles the Sun hangs low, so terrain throws long shadows that can hide rocks from a crew.

The model is given illumination angles, solar-frame anchors, and the tile footprint as sequence tokens, so it does not have to recover a quantity already stored with every tile. High-resolution and coarse tiles train in one mixed batch, and a FlexiViT patch embedding lets the same checkpoint move to other patch sizes without retraining the backbone. Pretraining ran on 16 H100 GPUs for 150,000 steps at a global batch of 1,536, about 1.1k GPU-hours, for a ViT-B encoder-decoder (768-wide, 12 layers, 12 heads).

Crater Detection Barely Beats the Baseline

NASA wanted three first uses: uncatalogued small craters, irregular mare patches that may be younger than the usual volcanic timeline, and polar ice. Scientists have already catalogued more than 2 million large craters. Polar rims that stay in sunlight are also candidate pads for solar power. Michael Barker, a NASA lunar topography expert who co-led the project with IBM, said the ages of irregular mare patches remain a matter of great debate, and that mapping their distribution is how that argument gets settled.

The published scores, means over five seeds, are less even than the ice headline.

DOWNSTREAM SCORES AGAINST THE BEST BASELINE

Task Metric NASA-IBM LFM Best baseline
WAC craters, 50% labels mAP (higher) 0.2541 (full fine-tune) 0.2313 (SwinV2-B)
WAC craters, full labels mAP (higher) 0.2581 (LoRA) 0.2420 (SwinV2-B)
NAC metre-scale craters mAP (higher) 0.1543 (LoRA) 0.1552 (SwinV2-B)
Irregular mare patches IoU (higher) 0.5709 (frozen) 0.5687 (ConvNeXtV2-B)
Polar ice prospectivity RMSE (lower) 0.0293 (full fine-tune) 0.0377 (SwinV2-B)

At metre scale the crater scores are low across the board, and the card says to treat the leaders as comparable because the margin sits inside the seed spread; part of that benchmark is annotated at 5 metres per pixel and looks blurrier. Irregular mare patch segmentation is the same story on the leaderboard, 0.5709 against 0.5687, but the random-init control collapses to 0.3142, so pretraining is doing real work there even if the baseline gap is tiny.

Label efficiency shows up on the wide-angle crater task. With half the training labels, the pretrained model already reached 0.2541 mAP, above SwinV2-B’s 0.2313 at that same 50% split and above SwinV2-B’s 0.2420 on the full set. LoRA, which IBM says left 90% of the base weights frozen, matched or beat full fine-tuning on crater detection. Full fine-tuning kept the edge on ice. NASA also showed the model outlining a new crater from a rocket-body impact near Einstein crater in an image held out of pretraining, next to older craters drawn in blue.

The Checkpoint Holds No Geodetic Frame

The second-order product is a shared encoder that a small lab can point at polar tiles. It is not a surveyor. The Hugging Face page lists the model card’s stated limits in language that is easy to skip under a 22% headline.

OUT OF SCOPE ON THE CARD

  • No geodetic frame: Generated latitude and longitude can be off by tens of degrees, and elevation can keep its shape at a shifted height.
  • Not measured ice: Polar outputs follow a fuzzy-overlay prospectivity map, not a detection of water in the ground.
  • Not a landing tool: The model is not validated for landing-site certification or hazard clearance.
  • Not calibrated generation: Any-to-any generated fields are qualitative probes, not instrument substitutes.

Those lines sit on the same page as the download button. Ablations that would isolate geometry tokens and mixed-resolution training from lunar pretraining as a whole have not been run, except for the random-init ice control. Some test sets are small. The frozen encoder, which won the 100-tile irregular mare patch split, fell below every baseline on crater detection, so a user still has to pick an adaptation method per task.

That is the gap between the announcement and the checkpoint. IBM framed the model as a rough guide for going back to the Moon. The card tells a downstream user not to certify a pad with it, and not to treat a yellow polar blob as a core sample.

The Weights Are Public on Hugging Face

The backbone, tokenizers, SomBench pretraining tiles, and four task benchmarks are public. Fine-tuning runs through IBM’s TerraTorch toolkit. NASA’s IMPACT team at Marshall Space Flight Center built the model with IBM Research and scientists at Goddard, Ames, the Universities Space Research Association, the SETI Institute, the University of Maryland, Baltimore County, and Howard University. Corresponding authors are Sujit Roy at NASA Marshall and Paolo Fraccaro at IBM Research.

IBM News posted the release the same morning.

The lunar model sits in a line of NASA-IBM science checkpoints, all open, each aimed at a different observing archive.

THE NASA-IBM MODEL FAMILY

  1. Early 2022: NASA and IBM begin foundation-model work under a Space Act Agreement, later folded into the Office of the Chief Science Data Officer’s AI-for-science push.
  2. August 3, 2023: Prithvi, a temporal vision transformer on Harmonized Landsat Sentinel-2 imagery, goes up on Hugging Face for flood, burn-scar, and crop tasks.
  3. September 23, 2024: Prithvi Weather-Climate, trained on MERRA-2, is released with a gravity-wave fine-tune.
  4. April 22, 2025: IBM and ESA open-source TerraMind, the multimodal Earth recipe the lunar team later retrained from scratch.
  5. September 10, 2026: The NASA-IBM Lunar Foundation Model and SomBench ship on Hugging Face, with fine-tuning code on GitHub.

A ViT-B trained in about 1.1k GPU-hours is small enough that a university group can LoRA-adapt it on a polar tile stack. That is the product NASA and IBM actually put in the open: a shared encoder for lunar tiles, strongest where it copies an ice-likelihood map, and explicit about the jobs it does not do.

Frequently Asked Questions

Where Can Researchers Get the NASA-IBM Lunar Foundation Model?

Weights and tokenizers live under the nasa-ibm-ai4science organisation on Hugging Face, and TerraTorch can build the backbone from the registry name nasa-ibm-ai4science/NASA-IBM-Lunar-Foundation-Model. Task configs in the NASA-IMPACT repo include configs/finetune/ice_prospectivity.yaml, so a lab can fine-tune the ice head without writing a new trainer.

Does the Model Detect Water Ice on the Moon?

No. The ice target is a knowledge-driven fuzzy-overlay prospectivity map, and a modality-count ablation found that the lunar model with only three layers (aspect, slope, and the DICE ice-stability field) already matched ConvNeXt-B using the full eight-layer stack, 0.0434 RMSE against 0.0437.

What License Covers the Weights and Data?

The model is Apache-2.0. SomBench pretraining tiles and the four task benchmarks are released CC BY 4.0 on Hugging Face, so the maps can be reused with attribution even when the backbone is swapped.

Can the Same Checkpoint Run at a Different Patch Size?

Yes. FlexiViT patch-embedding interpolation is built into the TerraTorch wrappers, and the published downstream numbers use a working patch size of 8 even though pretraining used 16-pixel patches on 256-by-256 tiles.

Who Led the NASA-IBM Lunar Foundation Model Work?

Sujit Roy (sujit.roy@nasa.gov) and Paolo Fraccaro (paolo.fraccaro@ibm.com) are the corresponding authors, NASA funded the work under Award No. 80MSFC25M0084, and Michael Barker at NASA Goddard co-led the science side with IBM.

Harry is the editor of TL TALK RADIO, an independent title he owns outright and edits himself, and much of his method comes down to one question: what was actually said? After ten years in journalism that began with reporting and led to editing, he treats the transcript, the recording and the written statement as the record, and a paraphrase from a third party as a lead to be checked, not a fact to be printed. Quotes on the site are matched to their source before they run. The same standard covers the whole publication, which serves readers across the world with news and sports, business and technology, science, entertainment, lifestyle, travel, auto and gaming. Figures are verified against the filing, dataset or scoreboard they came from, and mistakes are corrected on the page with a note saying what changed and when, as set out in the site's corrections policy. Readers who want to challenge a quote or a figure can write to support@tltalkradio.org.

Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending