Skip to content

The data gravity problem

CH 05 / 121,232 WORDST-06:00 READ

Eventually, moving all raw planetary data to Earth becomes irrational.

Capture first, transmit later, process on Earth. That sequence defined Earth observation for decades, and it treated the satellite as a camera that happened to be in orbit. The spacecraft sensed. Intelligence lived on the ground.

That architecture made sense when observation volumes were small. But the economics collapse once the planet becomes continuously observable across many sensing modalities at once: high-frequency imaging, hyperspectral sensing, orbital video, persistent SAR, thermal monitoring. Raw planetary data is growing faster than the infrastructure built to move it. Transmission becomes the bottleneck, and sending everything home stops being difficult and starts being irrational.

the cost curve flips. moving the data costs more than computing on it where it was born.

Most people outside the industry think the hard part of EO is launching satellites. It is not. Launch costs are falling, satellite manufacturing is standardizing, and sensors keep improving. The real constraint never shows up in a datasheet. Bandwidth.

the constraint nobody markets. it doesn’t appear in datasheets, but it bounds everything downstream.

A satellite can only transmit data to Earth when it has line-of-sight access to a ground station or relay network. Even then, transmission capacity is limited by physics, spectrum allocation, power budgets, antenna design, and orbital geometry. Meanwhile sensor generation rates keep climbing. The Sentinel-2 constellation produces about 1.6 terabytes of compressed image data a day, and its X-band downlink runs at 560 megabits per second. Those two numbers do not comfortably fit together, which is why ESA eventually routed the data through an optical laser link to geostationary relay satellites instead. The radio link had stopped being big enough.

Observation scales faster than downlink capacity, and every EO company eventually hits the same wall: you can’t transmit everything. So the question changes from “how do we collect more data” to “what data is worth transmitting”. That question is the beginning of orbital computation.

Historically, downlink systems were treated as infrastructure problems. Build more ground stations, lease more bandwidth, improve compression pipelines. But the scale transition underway changes the nature of the problem entirely. Imagine persistent global video coverage from orbit, or continuous hyperspectral monitoring across agricultural regions, or large-scale SAR constellations operating around the clock. The raw data volumes become staggering: petabytes per day, then exabytes.

[Artifact 05.01: Sensing outpaces downlink]

At that scale, transmitting raw sensor output stops being sustainable. Storage isn’t the expensive part. Movement is. Every transmitted bit consumes power, spectrum, orbital scheduling capacity, and antenna time. Hyperscale cloud systems learned this years ago. Once data movement dominates the cost, compute migrates toward where the data originates.

Earth-to-orbit communication still behaves more like early internet infrastructure than modern cloud networking: intermittent connectivity, limited throughput, high latency, strict scheduling. A satellite doesn’t hold a broadband link to Earth. It passes over communication windows.

This turns the spacecraft into a buffering system. Data accumulates onboard faster than it can always be transmitted. As sensing density increases, this imbalance worsens. More sensors create more observations, more observations create more backlog, and the pressure to filter rises with it. Eventually, the spacecraft must decide what should be prioritized, what should be compressed, what should be discarded, what requires immediate transmission. These are no longer transmission decisions. They are intelligence decisions.

the moment the satellite chooses what to send, it is already reasoning. the camera has become an editor.

Computing changed direction once engineers noticed storage was getting cheaper faster than bandwidth. Earth observation is at that same crossing. Orbital compute hardware is improving, onboard storage density keeps rising, and edge AI accelerators are becoming viable in space-qualified systems, while transmitting raw datasets to Earth stays constrained.

That inverts the economics. It may soon cost less to process data in orbit than to ship it home unread.

[Artifact 05.02: Pipeline inversion]

The satellite stops behaving like a remote camera and starts behaving like a distributed compute node. Onboard AI isn’t about faster inference. It is selective cognition. The system learns what matters before transmission happens.

inference at the sensor. the bit that never gets sent is the cheapest bit in the system.

Transmission delays are operational problems before they are bandwidth problems. Many future EO applications need near-real-time responsiveness: disaster monitoring, military awareness, maritime intelligence, infrastructure anomaly detection, climate event tracking, autonomous mission coordination. If sensing systems must wait for full downlink cycles before processing occurs, decision latency increases dramatically. In some operational contexts, the delay itself destroys the value of the information.

A spreading wildfire doesn’t care about orbital scheduling windows. A military maneuver doesn’t pause until imagery reaches a cloud region. A collision risk in orbit can’t wait for a centralized pipeline. So computation moves closer to observation. Autonomous vehicles can’t lean entirely on remote cloud reasoning, and neither can orbital systems. Local reasoning isn’t elegance. It is physics.

The future architecture of Earth observation looks a lot like distributed cloud computing, because the same forces are at work under different physics. Large-scale cloud systems evolved because centralized architectures stopped scaling. Running all computation through one location became slow, expensive, and fragile. So compute spread outward, closer to users and closer to where data was generated. Edge computing came out of that.

Earth observation is now replaying the same evolution at planetary scale. Satellites act as edge nodes, ground stations as regional ingress, orbital relays as network infrastructure, and constellations as distributed sensing clusters. The planet itself starts behaving like a continuously updating distributed system.

Eventually, entirely new architectural questions emerge. How do orbital systems coordinate state? How do satellites share learned representations? How do constellations synchronize models efficiently? How does inference occur collaboratively across fleets? These aren’t traditional aerospace problems. They are distributed systems problems appearing in orbit.

This transition changes the identity of satellites themselves. Historically, spacecraft were specialized hardware assets, carefully engineered and individually operated for one mission. But once onboard compute, continuous sensing, and distributed coordination become central, satellites begin behaving more like programmable infrastructure. Less like standalone machines. More like nodes in a planetary compute network.

Servers went the same way. A single server once mattered. Today infrastructure means orchestration across huge fleets. The future orbital operator may not think in satellites at all, but in observation capacity, inference coverage, temporal resolution, and reasoning throughput. The infrastructure layer becomes computational rather than aerospace.

the unit of work stops being “a satellite”. it becomes a slice of compute that happens to be in orbit.

Intelligence is the end product, not raw data. And intelligence doesn’t require transmitting every photon collected from orbit, only the abstractions that carry meaning. A future EO system may never send most of its raw observations home. It may transmit detected anomalies, infrastructure state changes, environmental alerts, object trajectories, and learned embeddings.

The output turns semantic rather than visual, the same way raw signals became structured understanding everywhere else in computing. Earth observation systems stop being imaging pipelines and become distributed planetary cognition running partly in orbit. At that point the line between space infrastructure and compute infrastructure stops meaning much.