Imagery is transitioning from product to infrastructure.
For most of the history of Earth observation, imagery was treated as a product. A satellite captured pixels. A company processed them. Someone purchased access. The transaction ended there. The whole industry organized around that assumption. Higher resolution meant better products, more revisit meant more valuable data, and better sensors meant a stronger moat. The dominant companies were the ones that owned satellites, controlled tasking systems, and distributed imagery catalogs. The pixel was the unit of value.
pick the unit of value and you have already picked the business model.
That model is collapsing. Imagery didn’t stop mattering; it became abundant, and abundance changes what a thing is for. The industry is moving from a world where pixels are scarce artifacts to one where pixels are continuous planetary infrastructure. The systems that matter in the next decade will be the ones that understand imagery, not the ones that collect it.
The original Earth observation stack optimized for acquisition. How do you launch better satellites? How do you improve revisit? How do you reduce cloud interference? How do you deliver imagery faster? These were rational questions for a world where obtaining imagery was difficult and expensive.
But as launch costs dropped, satellite manufacturing scaled, and constellations expanded, the economics started to invert. The bottleneck was no longer collecting pixels. The bottleneck became interpretation. A modern satellite constellation can generate more imagery in a day than entire organizations could process in a year a decade ago. Planetary-scale sensing is no longer the hard part. Extracting meaning from it is.
the early internet ran this same inversion. scarcity moved up the stack.
The internet did this first. Early on, access to information was the constraint. Then information became abundant, search emerged to organize it, recommendation systems emerged to personalize it, and AI systems emerged to understand and generate it. Earth observation is entering that same phase.
The planet has started producing data continuously, and raw imagery is worth little until something converts it into understanding. A 30-centimeter image of a port tells you almost nothing on its own. That vessel activity rose 17% week-over-week tells you something. That the rise breaks the port’s historical pattern tells you more. That the break lines up with a geopolitical event tells you most. The value is moving up the stack.
The most successful infrastructure eventually disappears. Developers rarely think about TCP/IP packets while building applications. Most users never think about databases while using software. Electrical grids are invisible until they fail. Imagery is moving in the same direction.
For decades, Earth observation products exposed pixels directly to end users because humans were still part of the interpretation loop. Analysts manually inspected scenes, compared imagery across time, and extracted observations. But machine learning systems are rapidly compressing this loop. Foundation models trained on planetary-scale imagery are beginning to treat satellite data not as photographs, but as machine-readable representations of the physical world.
Imagery was built for human eyes. Now it is being built for machine cognition. The consumer of satellite imagery is no longer an analyst. It is another model.
once a model is the reader, every format designed for human eyes is dead weight.
A logistics platform — doesn’t want pixels. it wants supply chain intelligence.
An insurance company — doesn’t want imagery. it wants flood risk estimation.
An agriculture platform — doesn’t want multispectral bands. it wants crop stress detection.
Pixels become intermediate infrastructure inside larger computational systems. Eventually, many end users may never directly interact with satellite imagery at all. They will interact with decisions, predictions, alerts, and simulations generated from planetary sensing systems operating continuously in the background.
Large language models changed software because they created a generalized representation layer for language. Instead of building separate models for summarization, translation, classification, extraction, and question answering, the industry discovered that a sufficiently large foundational model could internalize broad structures of language itself. Earth observation is moving toward a similar abstraction.
Foundation geo models are trying to learn the latent structure of the physical planet: relationships across geography, time, climate, infrastructure, and human activity, rather than objects in images. A traditional remote sensing pipeline acquires imagery, processes it, trains a task-specific model, detects an object, and delivers an output. Foundation geo models collapse large parts of this stack into shared planetary representations. Instead of training separate systems for roads, ports, construction sites, forests, or shipping activity, the model learns generalized spatial intelligence.
The physical world is deeply interconnected. Ports influence roads. Roads influence urban expansion. Urban expansion influences energy usage. Energy usage influences emissions. Emissions influence climate behavior. The planet is a coupled system, not a collection of isolated datasets. The next generation of Earth intelligence systems will model those couplings directly.
task-specific models treat the planet as independent slices. the planet doesn’t behave that way.
Raw imagery is information-dense but meaning-poor. Two satellite images separated by six months may contain millions of changed pixels, and a tiny fraction of them matter. A new warehouse matters. Seasonal vegetation shifts may not. Cloud movement rarely does. Military asset movement can decide a policy. The challenge isn’t seeing changes. It is ranking them.
Semantic understanding is what closes that gap, because semantic systems work at the level of concepts rather than pixels. Instead of asking what changed in this image?, they ask what changed in the real world? One word, and the architecture of the whole stack changes.
Traditional computer vision focused on detection and segmentation because the industry was image-centric. The future stack is world-centric. It has to reason about infrastructure growth, economic activity, environmental stress, supply chain behavior, and climate-driven change. That takes reasoning layers far above image classification. The image becomes evidence, not the product.
swap “world” for “image” and the pixel quietly drops from product to artifact.
Most Earth observation systems still think spatially before they think temporally. They answer what exists here? The better question is usually how is this changing? The planet doesn’t hold still. Cities expand, rivers shift, forests disappear, factories switch on, shipping lanes fluctuate, and conflict reshapes infrastructure. A single image is a snapshot. Planetary intelligence comes from sequences.
Temporal reasoning is becoming the defining capability of modern geo AI. Earth observation starts to look more like video understanding than static image analysis. The challenge is modeling behavior across geography and time, not detecting objects inside single scenes.
the unit of analysis stops being the scene. it becomes the trajectory.
What is normal behavior for this port?
What is anomalous activity for this refinery?
How does this region evolve seasonally?
What patterns precede drought conditions?
What signals historically appeared before supply disruptions?
That is planetary-scale behavioral modeling, well past computer vision. And once systems track planetary behavior continuously, Earth observation stops being a map layer and starts being a real-time intelligence system.
For years, image interpretation itself acted as a moat. Teams with better analysts, proprietary labeling pipelines, and domain-specific models could create differentiated value from the same imagery sources. But foundation models are rapidly compressing these advantages. Generalized vision systems are becoming dramatically better at extracting structured meaning from imagery with far less task-specific tuning.
That is dangerous for traditional EO companies. If everyone has similar imagery and similar interpretation models, then image interpretation alone stops being defensible. The value moves toward proprietary workflows, domain expertise, temporal datasets, feedback loops, and distribution.
databases, cloud compute, basic ML. every commoditized primitive took this same road. EO is next.
This mirrors what happened in software infrastructure. Databases became commoditized. Cloud compute became abstracted. Basic machine learning became accessible. The winners were rarely the ones holding the raw primitives. They were the ones who built the most useful systems on top. Earth observation is heading toward the same outcome.
Large language models taught Earth observation a lesson it has not absorbed yet. Before GPT-style systems, most NLP was narrow pipelines solving isolated tasks. Then scale changed the game. Once models got large enough and read broadly enough, capabilities nobody designed showed up on their own: reasoning, generalization, transfer.
Earth observation may hit the same discontinuity. Much of the industry still thinks in narrow workflows: detect ships, count vehicles, segment buildings, classify crops. Planetary-scale multimodal models may learn far richer abstractions about how the physical world behaves, not because anyone programmed them to, but because the underlying data contains latent structure waiting to emerge at scale.
nobody designed reasoning into GPT. it fell out of scale. the planet has more latent structure than language does.
This is why the future of Earth observation likely belongs less to imagery providers and more to intelligence layer builders. The companies that matter most may not be the ones launching the largest constellations. They may be the ones building the cognitive systems that understand the planet continuously.
The industry is slowly moving toward a new architecture. At the bottom sits sensing infrastructure: satellites, drones, airborne systems, IoT networks. Above that sits planetary data infrastructure for storage, indexing, processing, orchestration, and retrieval. Then comes representation infrastructure built on foundation geo models and multimodal embeddings. Above that, reasoning infrastructure handles temporal analysis, simulation, forecasting, and anomaly detection. At the top sits decision infrastructure: applications, agents, operational systems, automation.
Most of the original EO industry concentrated on the bottom layer. The top is where the returns are. Once imagery becomes infrastructure, the defining question is no longer who owns the pixels? It becomes who understands the planet best.
ownership was a legal question. understanding is an engineering one.
That question has no clean answer yet. The pixel layer is being built fast. The understanding layer is barely started. The next chapters trace what it would take to build it, and who is positioned to.