The Future of Remote Sensing: AI, Foundation Models and the Next Generation of Earth Observation

ESA's Φsat-2 is a miniature Earth-observation satellite designed to demonstrate how artificial intelligence can process observations in orbit. Source: ESA – The Φsat-2 satellite Credit: ESA

Earth observation has a data problem.

That may sound surprising when one of the great achievements of the remote-sensing community has been the enormous expansion of freely available satellite imagery.

But the challenge is no longer simply collecting observations. It is extracting useful information from them.

Every day, Earth-observing satellites produce vast quantities of imagery and other measurements. Analysing all of that data manually is impossible.

Artificial intelligence is therefore becoming one of the most important technologies shaping the future of Earth observation.

From conventional machine learning to foundation models

Machine learning is already well established in remote sensing.

Models can classify land cover, detect objects, map changes and estimate environmental properties.

Traditional approaches, however, often require carefully labelled training datasets for each task.

Foundation models offer a different approach.

Rather than training a model exclusively for one task, foundation models are generally pre-trained on very large and diverse datasets and then adapted to multiple downstream applications.

This concept has rapidly entered Earth observation.

Recent surveys identify remote-sensing foundation models as an emerging research area spanning visual, language and multimodal models, while highlighting the distinctive challenges created by different sensors, spatial resolutions and temporal dynamics.

Why Earth observation is different

It would be tempting to assume that a model developed for ordinary photographs can simply be applied to satellite imagery.

It cannot.

Earth observation data has its own complexities.

A satellite image might contain information from visible, infrared or microwave wavelengths. Images can have very different spatial resolutions. Radar imagery has fundamentally different characteristics from optical imagery. The same location can look dramatically different depending on season, viewing geometry, atmospheric conditions and sensor.

And geography matters.

A model trained on one part of the world may not perform equally well somewhere else.

This makes generalisation one of the central challenges for EO foundation models.

Multimodal Earth observation

One of the most promising developments is the move towards multimodal models.

Instead of considering satellite imagery in isolation, researchers are exploring systems that can combine optical imagery, radar, elevation, land-cover information, geographical context and other datasets.

ESA and IBM’s TerraMind provides an example of this direction. Released in 2025, TerraMind was developed as an Earth-observation foundation model capable of working with multiple types of geospatial information, including Sentinel-1 radar and Sentinel-2 optical data, alongside topography, elevation and land-cover information.

The attraction is straightforward.

The Earth is not a collection of independent pixels. It is a connected physical system.

A forest has a location, elevation, structure, climate, seasonal cycle and spectral response. A building has geometry, materials, surrounding infrastructure and a particular spatial context.

Models that can use several types of information may therefore be able to build richer representations of the world.

What could these models actually do?

Potential applications are remarkably broad.

ESA projects are exploring foundation models for tasks including disaster analysis, methane-leak detection, forest biomass monitoring, soil-property estimation, land-cover change and monitoring mining expansion into farmland.

Other initiatives are examining applications such as snow monitoring, flood mapping, sea-ice mapping and drought monitoring.

The significance is that a single underlying model could potentially support multiple applications rather than requiring a completely new model for every problem.

That could make EO analysis considerably more scalable.

The emergence of Earth-observation “embeddings”

Another interesting development is the use of embeddings.

An embedding is a numerical representation that captures information about an input in a form that can be used for further analysis.

In June 2026, ESA reported the wider availability of Tessera, a foundation model trained on Sentinel-1 and Sentinel-2 observations. The model generates representations of what satellites observe that can be used to create information-rich maps and support applications including crop monitoring, burned-area mapping and forest-canopy analysis.

This illustrates a potentially important change in how we work with EO archives.

Instead of repeatedly processing every pixel from scratch for every application, researchers may increasingly work with reusable representations derived from large-scale pre-training.

AI doesn’t eliminate the need for remote-sensing expertise

This is perhaps the most important point.

AI does not make remote-sensing knowledge obsolete.

In fact, it may make that knowledge more important.

A model can identify patterns without necessarily understanding the physical process behind them.

A classifier might associate a particular spectral pattern with flooding because it has seen similar examples. But if atmospheric conditions, geography or sensor characteristics change, the model might fail.

This is known as a form of distribution shift, and it is a major concern for real-world deployment.

Recent research into EO foundation models continues to identify issues including multimodal data alignment, scalability, training datasets and transfer between different applications.

Accuracy isn’t the only measure

The remote-sensing community has traditionally placed strong emphasis on accuracy assessment.

That remains essential, but foundation models introduce additional questions:

Can we explain why a model made a particular prediction?

Does it behave consistently across different geographical regions?

How does it respond to a sensor it has not encountered during training?

Does high benchmark performance translate into reliable operational performance?

And what happens when the model encounters an unusual event?

These questions matter particularly when EO-derived information influences decisions involving environmental management, disaster response or infrastructure.

AI at the satellite

Another emerging direction is moving some processing closer to the sensor.

Rather than transmitting every observation to Earth and processing everything on the ground, satellites can potentially use onboard AI to identify important information before transmission.

ESA research is exploring onboard machine learning for time-critical applications including methane detection, vessel detection, fire detection and flood detection.

This could be significant because satellite communications and ground-processing capacity are finite.

If a satellite can identify potentially important events in orbit, it may be possible to prioritise what is transmitted.

For rapidly developing events, that could reduce the time between observation and response.

The environmental cost of AI

There is another question that deserves attention: sustainability.

Training very large models requires substantial computational resources. As EO models become larger and datasets become more complex, the environmental and financial costs of training and deploying them need to be considered.

ESA’s Φ-lab has already begun exploring more sustainable approaches to training Earth-observation foundation models, reflecting a growing recognition that AI innovation itself needs to be evaluated in terms of resource consumption.

For the EO community, this creates an interesting tension.

We are using increasingly sophisticated technology to understand environmental change, while needing to consider the environmental cost of the technology itself.

The role of the geospatial professional

The future is unlikely to belong solely to AI specialists or solely to traditional remote-sensing experts.

Instead, it will favour people who can work across disciplines.

Understanding sensor physics, coordinate systems, photogrammetry, image processing and uncertainty remains valuable. So does understanding machine learning, cloud computing and data engineering.

The strongest EO practitioners may increasingly be those who can connect these worlds.

A new relationship with Earth observation data

For decades, remote sensing has involved a workflow in which experts formulate a question, select imagery, process it and derive a result.

Foundation models may change that workflow.

Instead of beginning with a single task-specific algorithm, analysts may increasingly begin with a large pre-trained geospatial model and adapt it to a particular question.

That does not mean the traditional workflow disappears.

It means the balance between data collection, processing, modelling and interpretation may change.

The technology is still developing, and there are genuine limitations. But the direction is clear: Earth observation is moving towards increasingly automated, multimodal and intelligent analysis.

The challenge for the remote-sensing and photogrammetry community is to ensure that this transformation is not simply technologically impressive, but scientifically sound.

The future of EO will not be defined by whether AI can look at an image.

It will be defined by whether AI can help us make better, more reliable and more scientifically defensible observations about our planet.

Sources

Discover more from RSPSoc.org.uk

Subscribe now to keep reading and get access to the full archive.

Continue reading