News22 min read

Introducing Sonic Inference Pods: Modular data centers built for inference

Sonic Inference Pods are Runware's modular 1MW inference data centers: 1,200 GPUs in a 20-foot container, closed-loop liquid cooling, and 30–80% lower cost per GPU-hour. Here's how we got here.

A Runware Sonic Inference Pod: a modular shipping-container data center built for AI inference
Ioana Hreninciuc
Ioana Hreninciuc

Today we're announcing Sonic Inference Pods, our modular data centers built for inference. Some of you will know them as Sonic Inference Pods, which is what we called them internally for most of their life.

A Sonic Inference Pod is 1MW of IT compute in a 20-foot shipping container, with a chiller roughly the size of the container sitting on top of it. Inside: around 1,200 GPUs across custom servers with no cases, custom racks, our own PCIe switching, high-frequency CPUs, local NVMe, and a closed-loop liquid cooling system that consumes no water. It arrives on a truck. It needs ground, power, and a network connection. It runs as one node in a single distributed inference network.

Pods are deployed in the United States and Europe. Over the second half of 2026 we will bring up to 10,000 nodes online, and we're targeting more than 1GW of inference compute in 2027. Because we design, build, and operate the hardware ourselves, inference on Sonic Inference Pods costs 30–80% less per GPU-hour than other inference providers depending on the workload. For most workloads it's 50% or more. You can run your models on that capacity today through Runware Serverless.

That's the announcement. The rest of this post is how we got here, and I'd argue it's the more useful part. We did not set out to build data centers. We set out to make inference fast enough and cheap enough to build a product on, and after three years of chasing that, the data center was the last place the cost was hiding.

We do inference. We don't sell data center space.

Worth saying up front, because "startup builds data center" invites the wrong comparison. We're not trying to out-build a hyperscaler, and we don't sell square footage, racks, or colocation. We sell inference. The building is a cost input, not the product.

We build the building because of where the money goes. Sarah Friar, OpenAI's CFO, has put the cost of a gigawatt-scale AI data center at roughly $15 billion for the facility alone, before a single GPU goes in it. Every inference provider renting space in someone else's building pays a share of that, plus the operator's margin, plus the cost of cooling GPUs inside a facility designed for web servers. That was the one layer of our stack we hadn't touched. So we took it out, and what we save there we pass on in the price.

Here's how we got to that conclusion, which took about four years and several things going wrong.

2022: a product that couldn't afford to exist

Stable Diffusion's weights went public in August 2022, and 1.5 followed in October. Like a lot of people, my co-founder Flaviu Radulescu and I thought it was magical technology. It was also extremely slow. Beyond everyone having seven fingers, generating a handful of 512x512 images took close to a minute, sometimes longer.

Flaviu had an idea that sounded implausible at the time: you should be able to use image generation the way you use Google Images. Type a prompt, scroll the results, watch them appear in real time. If that were actually possible, surely one of the labs would already have done it.

He built it anyway. He called it PicFinder, the tagline was "image generation as fast as image search," and he put it on the internet with no plan whatsoever. Some YouTubers found it and featured it. It went from nothing to millions of users in less than two months, and passed 100 million images in three. Generation took under a second while comparable apps took thirty.

He also stopped sleeping, because keeping the GPUs and the product up was a full-time job on its own. I had another job at the time, so what I could usefully do was help pay for the GPUs. We started talking about whether there was a company here. I applied to a16z Speedrun, we got in, I quit my job, and in 2023 we went to San Francisco and got our first check.

Then we came back to London and ran into the wall. The GPU prices we could get made the product impossible. We wanted people to generate effectively unlimited images; the cost per image wouldn't come down far enough to support that. This is the first time we talked seriously about owning hardware.

The GPUs we rented were not the same as each other

We were buying capacity wherever it was cheapest: marketplaces, aggregators, spot instances, whatever we could find. And the quality varied enormously. Same GPU, same price, meaningfully different throughput.

We started testing to work out why, because we needed to understand what a low enough cost per image would take. The answer was that the parts around the GPU mattered far more than anyone was accounting for. CPU clock speed in particular. Nobody selling us GPU hours was optimizing for that, because they weren't selling inference. They were selling GPUs by the hour, and those are different products.

That was the moment hardware stopped being a cost question and became a control question. You can't tune what you don't own.

Summer 2023: three servers on a kitchen table

We had both run infrastructure before and managed clouds, so this felt less mad to us than it probably should have. We bought three nodes with liquid-cooled GPUs, three GPUs each, mostly as an experiment to see what performance we could actually get. They were delivered to my house.

We put the first server case on my living room table, which was also my kitchen table. We decided to document the build as a step-by-step guide, on the reasoning that if we bought more of these, Flaviu was not going to be the person assembling all of them. We assumed the first one would take an afternoon.

It took three days. Not three twelve-hour days, three to six hours each, but three days. The liquid cooling was genuinely intricate: dozens of small parts, everything needing to fit in a particular order. The documenting wasn't what slowed us down; there were two of us. It was just harder than it looked. Every evening we said we'd finish tomorrow. On the third day we finished because we'd run out of tomorrows.

That server then ran in my home office for the entire summer, producing an unbearable amount of heat in a London flat with no air conditioning. It also worked, well enough that we kept it running and used it to test things. We spent the summer swapping components, different CPUs, different RAM, trying to find what actually moved the numbers. Between the hardware changes and the software work happening alongside them, we got roughly 50% more performance out of the same GPUs.

The pivot we didn't want

While that was going on, we tried to raise money for the image generator and failed. In the process we worked out something more useful than the money would have been: we didn't have founder-product fit for a consumer app. Both of us had spent our careers building deeply technical platforms. Neither of us was the right person to run an image generator.

What we did have was the first real-time inference engine for media. So we stopped trying to be the image generator and started building the engine underneath all of them. We called it Runware, went back out to raise, and this time it worked. An oversubscribed pre-seed closed in December 2023.

2024: nobody would take our servers

By then we'd gone from three custom servers to a few dozen, and we were colocating them in data centers. That turned out to be its own problem.

GPU servers draw far more power than the servers those facilities budget for, even with the smaller GPUs we were running then. Operators priced accordingly. Worse, almost none of them supported liquid cooling. In 2023 and 2024 it was very hard to find any data center that would take a liquid-cooled rack at all. So we air-cooled GPUs, which is expensive, inefficient, and throws away a good part of the efficiency we'd just spent a summer engineering in.

This is where it stopped being a procurement annoyance and became the thing we believed. We think every product will eventually touch a GPU. All of those GPUs need somewhere to go, and the data centers that exist don't have the power or the cooling to take them at scale. That isn't a shortage that clears. It's a mismatch between the buildings that exist and the workload arriving.

So at the end of 2023 we started designing our own: the servers and the data center around them, together.

What we built, and why each part is the shape it is

The software platform had to exist first. Most of the first half of 2024 went into building Runware itself, which launched in October 2024 and scaled quickly. Our GPU spend went up with it, which settled any remaining argument about whether we needed our own hardware.

Hiring for it was easier than we expected. Hardware engineers aren't in the bidding war software engineers are in, and we hired people we'd worked with before, so we had a team by the beginning of 2025. Flaviu had been head of R&D and managing director at an infrastructure company in the early 2000s, running data centers with hundreds of thousands of servers, and later co-founded a bare-metal cloud. He'd designed hardware before. He conceived the pod and led the engineering, and that's why the first concept and the finished thing look so similar. We knew which problems we were solving before we started drawing.

Servers with no cases

Our servers don't have cases. They're shelves that click into place like Legos. That decision comes directly from three days at my kitchen table. It isn't enough for a data center to be efficient to run. It has to be efficient to install. We couldn't have people assembling servers for weeks. The cooling is already routed into the racks so everything lands where it's supposed to. Installation takes hours.

Our own PCIe switch

Having established that CPU frequency matters more than anyone was pricing in, we chose high-frequency CPUs that clock up to 6GHz. Those don't carry enough PCIe lanes for the GPU density we wanted. So Flaviu designed our own PCIe switch. GPUs connect to the CPU and to each other through it. Nodes run 2 to 8 GPUs depending on the workload, and the architecture supports up to 16.

Local NVMe, so models stay warm

One of the real limits on running hundreds of thousands of models is loading time. When you bring up more servers, you have to get the weights onto them, and pulling hundreds of gigabytes from distributed storage is slow. Each of our servers carries enough local NVMe to hold the current generation of large models, frontier LLMs, video models, image models, on the node itself. The whole path is optimized for loading weights into GPU memory as fast as possible, over PCIe. Models are always warm and always local, so cold starts largely stop being a problem.

Cooling nobody would sell us

There was no off-the-shelf system that could cool a megawatt at the density we wanted. We looked at several, worked with a number of companies on a design, and it took an enormous amount of input from Flaviu. The first cooling system we commissioned was never delivered, and we had to start over with someone else. We ended up sourcing components from several suppliers and built it ourselves.

Along the way, expert cooling engineers advised us to fill the system with the wrong liquid. He caught it. That's the difference between hardware and software: there's no debugger, mistakes take months rather than minutes, so you validate everything yourself in more detail than feels reasonable.

What came out of it is a single closed loop that recirculates about 1.5 cubic meters of liquid and holds GPU temperatures within 2°C of target in ambient conditions up to 50°C. Nothing is plumbed in. It has no connection to water mains and consumes no water in normal operation. Physically, there is nothing left to make more efficient.

A shipping container, and a crane

All of it had to fit in a 20-foot container: three rows of servers, one along each side wall and one down the middle, with two aisles to walk between them. The center row is single-sided, with the back of it carrying the rest of the equipment. That comes to around 1,200 GPUs in a shipping container. The chiller sits on top and is about the same size as the container underneath it.

"It ships on a truck" makes this sound simpler than it is. Moving a pod means a crane, which means a crane operator. Then you position the container, lift the chiller onto it, and connect the two. We've now done that enough times to know exactly how long each step takes, which is a different kind of knowledge from designing it.

Any GPU, including the next one

The first working prototype went into production in early 2025 with smaller GPUs, because we still had Stable Diffusion workloads and we were being careful about the most expensive component. It became clear quickly that this wasn't enough. Video models were much larger, and we knew we'd want to run LLMs. So we upgraded that first pod, with the additional funding, to handle any class of model.

The cooling blocks fit multiple GPU types. We run RTX PRO 6000 as the workhorse node because we tested what actually delivers throughput per dollar, and we can also run B200 and B300. We expect to be among the first providers deploying Vera Rubin GPUs at scale, because when a new generation arrives we update the pod, not the building.

From mid-2025 to now, the work was optimization: getting the PCIe switch, the components, and the interconnects running at the speeds they should. Much like shipping software and then tuning it, except each iteration takes months. Once the pod was fully optimized, we ran our own inference on it.

We built it for ourselves. Then the market ran out of room.

At the start of 2026 the ground moved. Data center capacity started running short, announced large data center projects were cancelled or deferred, and local moratoria on new data center construction began appearing in places that had been building without much friction. Everyone needed somewhere to put GPUs, and the places to put them were getting harder to build, not easier.

We already had the answer, because we had spent three years building it for our own product. So we ramped production.

This is the second time we've done the same thing. We built a real-time inference engine because our image generator needed one, then made it an inference engine anyone could use. We built modular data centers because our inference engine needed somewhere to run, and now we're making them available to everyone else.

That's why we're deploying the first 10,000 nodes, why we have our first 160 locations booked, and why we're currently bringing modular data center capacity online faster than the rest of the world combined. There'll be more news on this shortly.

Why inference on Sonic InferencePods costs less

Every saving traces back to something above.

  • No facility to amortize. A pod needs ground, power, and connectivity, not a hundred-acre engineered building.
  • No overhead we don't use. No raised floors, no oversized redundancy tiers, none of the enterprise data center apparatus that general-purpose workloads require and inference doesn't.
  • No inefficient cooling. Liquid, closed loop, designed for this density instead of adapted to it.
  • No unnecessary components. No server cases. Custom PDUs. Power supplies bought direct from wholesale vendors. Our servers are built for inference and nothing else, so we don't pay for anything else.
  • Single tenant facilities. A pod site runs our workloads only, so there is no shared-facility margin to pay. (Serverless capacity is multi-tenant at the software layer; dedicated and bare-metal pods give you a hardware boundary.)
  • Manufactured, not constructed. Pods are produced and shipped. No permitting queue, no multi-year build, no grid interconnect wait.
  • Power bought at generation price. In our benchmarking, power is around 30% of the lifetime cost of running a GPU. We put pods at wind farms, solar parks, and near hydro, buying at generation price rather than grid price, with no transmission losses between the generator and the GPU. This is the largest single lever.

At the facilities and deployment layer, our pods cost roughly $150M per gigawatt against about $15B on the traditional path, up to 100x lower, with GPU hardware excluded on both sides. Which comes out, at the level that matters to you, as 30–80% lower cost per GPU-hour than other inference providers, and 50% or more for most workloads. Where a given workload lands in that range depends mostly on the model: its size, how it parallelizes, and how well it fits the node.

Where we can put them

We have 160 locations available to us today: sites with power where we can place containers now. Cooling is redundant. Networking is redundant. Where a site has redundant power we use it; where it doesn't, we can trade that for very cheap power, because pods are built around similar workload profiles and traffic reroutes across the network instead of relying on local redundancy. The blast radius is one pod, not one data center.

Because a pod is a shipping container, it can be deployed inside any national border or jurisdiction. Organizations with data residency or regulatory requirements can run inference on dedicated, locally sited infrastructure without waiting years for in-country construction. We say internally that we're the only provider who could put a data center in the Vatican. If a customer wants one near their office, or on their roof, that's a logistics conversation, not a construction project.

The other reason sites are available to us is that pods don't ask anything of the local utilities. A growing share of data center projects are now blocked or delayed because they can't secure power and water, and because communities object to a facility drawing millions of gallons for cooling. Our cooling loop is closed. Pods need no water mains and consume no water in normal operation, there is no evaporative cooling, and there is no hundred-acre site to clear. They can run entirely on renewable power. Where moratoria on new data center construction are becoming a real constraint on capacity, essentially none of the objections apply to a container sitting on an existing power site.

Deykhan Ten, VP of Strategic Partnerships at Higgsfield, put it better than we would:

"We started with Runware's Model API, but quickly expanded into their inference infrastructure because of the scale and efficiency they could deliver. Our models reach millions of users every week, so reliable capacity and cost-efficient inference are critical. The lower our inference costs, the more value we can pass on to our users. Runware understands that deeply, which is why we work closely with them on our most important model launches."

Capacity beyond our own pods

The pod fleet is the engine, not the ceiling. We manage inference capacity across both our own pods and hyperscaler networks, so when you need more capacity, or need it earlier, or need it in a region where we haven't landed a pod yet, we extend into those networks rather than telling you to wait. Same platform, same contract, same routing layer.

The cost advantage lives in the pods, which is why we keep building them. But no workload should be capped by how fast we can ship containers.

What happens next

Our $50M Series A in December 2025 is what makes the rollout possible. Over H2 2026 we will bring up to 10,000 nodes online across the United States and Europe. We're targeting more than 1GW of inference compute in our pods in 2027, and we'll say more soon about how we plan to get there.

We'll also write in more detail about the engineering: the cooling design, the PCIe switching, the distributed network and routing, and the decisions we'd make differently.

Running your models on Sonic Inference Pods

Capacity is available now through Runware Serverless. Deploy your own models, containers, and AI workloads while we handle provisioning, scaling, and operations. What we commit to is the economics rather than any particular price: we intend to offer the best cost per inference in the industry, and owning the hardware is what lets us keep doing that as the hardware changes.

If you'd rather see the numbers before you commit, we'd prefer that too. We'll benchmark your model as-is on the exact hardware that would serve it, profile it across the GPU options that fit, and work with your engineers on the serving path. You get a clear view of your unit economics before you decide anything. We'd rather show you results on your model than publish numbers from ours.

We're talking to frontier labs, AI studios, enterprises, and organizations with sovereign or data residency requirements.

The goal hasn't changed since the server on the kitchen table. We want inference to cost as close to the price of electricity as it possibly can. Owning the building, and putting the building where the electricity is, is how we get there.

Articles|

Run the fastest, lowest-cost generative AI API.

Start with free test credits.

Get started now