The Edge Evolution: AI everywhere
Our new paper "The Edge Evolution: AI everywhere" covers the development of edge compute from IoT and smart sensors through cloud gaming to distributed AI inference, the four edge layers and what belongs in each, the sovereignty and privacy case, the evolution of MEC and split compute, and what the next phase of edge buildout requires.
Inference is moving to the edge
Most of the capital and most of the coverage in AI infrastructure has thus far gone to training. That is already changing. Gartner puts AI-optimized IaaS spending at $42.3 billion in 2026, with $23.3 billion of it going to inference, against $19 billion for training, and expects inference to account for 59 percent of the total by 2027.
Training builds the model. Inference is what the model does for a living. It runs every time a customer asks a question, every time a sensor reports a reading, and every time a camera frame needs a decision made about it. It is the production face of AI, the real-world intelligence, and it behaves nothing like training.
Our new paper, "The Edge Evolution: AI everywhere", examines what that difference means for the physical locations where compute gets built.
Different workload, different geography
Hyperscale is not going anywhere. Goldman Sachs expects the five highest spending hyperscalers to invest a combined $400 billion in such facilities in 2026, and large model training will keep absorbing it. Centralized capacity is the right shape for a workload that runs for weeks, tolerates distance, and cares only about raw density.
Inference is the opposite case. It runs constantly, in short bursts, and it is judged on response time. A query that travels several hundred miles to a centralized facility and back has spent most of its budget on the journey. For a chatbot that is an irritation. For traffic management, fire detection or a robot on a production line, it is the difference between a system that works and one that does not.
So the compute follows the data. That is the argument the paper makes, and it is why edge infrastructure has stopped being a promising idea and started being a build requirement.
The edge is not one place
The paper sets out four distinct layers, and the distinction between them matters physically and commercially.
1. The device edge is where data originates: a phone, a doorbell, a smoke sensor.
2. The on premises edge sits at a factory or utility plant, processing what the site generates.
3. The network edge covers small compute installations designed for specific workloads, deployed at cell towers and network hubs.
4. Distributed data centers are the AI compute hubs embedded inside telecom networks.
Each step closer to the source cuts latency. Which layer a workload belongs in depends on what it is doing, not on architectural preference.
Targeted inference beats general capacity
An edge node does not need to serve every possible workload. It needs to serve the ones its location generates, which makes it a smaller and far more predictable build than a hyperscale facility.
A manufacturer can deploy a node sized for the sensor data its own plant produces, analyzing productivity and flagging mechanical failure before it happens. A local authority can run a small language model trained on the queries its residents actually ask, covering transport, tourism and local services, and refine it as usage teaches it what people want. A national government can prioritize live translation based on who is visiting and in which languages.
None of these need frontier scale compute. They need compute that is close, always available, and specific.
Sovereignty is a practical consequence
The regulatory case arrives at the same answer as the performance case. Gartner expects more than 75 percent of European and Middle Eastern enterprises to geopatriate workloads by 2030, up from under 5 percent in 2025.
When inference happens at an edge node a few miles from the user, the data has not crossed a border, and the question of which jurisdiction it landed in does not arise. For developers building AI applications under data residency obligations, processing locally and returning only the result is the cleanest available answer. The paper covers this in detail, including how split compute divides a workload between a capable device and a network edge layer, and keeps sensitive data on the device entirely.
What the hardware has to survive
Edge sites are not data halls. Space is limited, power is limited, and access can be difficult. Cell towers are built where coverage demands it, not where engineers can reach them easily.
That puts two requirements on any edge AI installation: get as much compute as possible into the available footprint, and keep it running without regular intervention. Immersion liquid cooling addresses both. It supports high density in constrained space, protects hardware in hostile environments, and its modular pod format scales from a single pod at a tower to multi pod deployments in exchanges and telco data centers.
Operators hold the ground
Telecom operators already own the two hardest components of an edge network: physical locations at the right distance from users, and the high speed connectivity linking them. Nineteen operators recognized this in 2020 when the Telco Edge Cloud Initiative launched, and the GSMA now describes AI as adding a new dimension to the value of edge through distributed inference.
Edge AI capacity is a route for operators to move up the value chain, using assets they have already paid for. The paper works through what that build looks like in practice.
Who should read this paper?
It is written for telecom operators working out what AI inference means for the network assets they already own, and for enterprise and public sector teams whose workloads cannot afford the round trip to a centralized facility. Compliance and security leads will find the data residency case in section 4 stands on its own. Developers architecting inference across device and network layers should start with the split compute material.

.avif)