Skip to main content
FIELD STATION· an action lab
EXPERIMENT· ENTRY 57

Shared inference: when several ordinary computers become one AI system

Eight people, eight computers, one model. It worked, until somebody closed a browser tab.

Tom Watson · 30 Sep 2026 ·experiment, sovereignty, society
# experiment# sovereignty# society
Several people's hands spreading out a large map pieced together from separate sheets
FIG 1 Image credit: FIELD STATION

Eight people, eight computers, one model.

For a few minutes during our online session, Qwen 27B was running across the laptops and desktops of the people on the call. No data centre. No single machine doing the thinking.

It worked, but then somebody closed a browser tab, and everything stopped.

That moment just about sums up our first FIELD STATION experiment better than anything else. We proved several ordinary computers can run an AI model together. Keeping it running while those computers come and go is a different problem, and a harder one.

Where we started

The experiment began with a fairly simple question: could a group of people run AI for one another using hardware they already owned?

Our original plan was to test this using co/core. A small group would act as both users and providers: sometimes asking somebody else's machine to run an AI job, sometimes contributing their own.

Completed jobs would leave a signed public record, and the prompts and outputs would stay private. It was partly a technical experiment, but also an experiment in reciprocity. What might AI look like if compute were treated as shared infrastructure rather than something rented from a distant platform?

Then we found another way of asking the same question...instead of passing complete jobs between different people's computers, what if the computers could run the same model together?

From SwarmLLM to Shared AI

That became Shared AI, built on the open-source SwarmLLM. We explained the switch in our revised experiment post.

Rather than every computer needing to hold a complete model, different machines hold different parts of it. Each browser contributes memory and GPU capacity. The model's layers are divided between the participating devices, and the activations (the intermediate results passed from one layer to the next) move directly between them using WebRTC.

No central data centre, just several ordinary computers, temporarily becoming one AI computer.

Other people are asking this too

We're not the only ones exploring this. It helps to separate two questions: who owns and governs the system, and where does the computing happen?

Darkbloom, from the a16z-backed Eigen Labs, spreads the computing out. It routes encrypted1 requests to people's Macs and pays the owners for their time. Ownership and control stay with a venture-funded company, which is not a direction we want to go in.

The Inference Cooperative answers the other question. Its members govern it, one member, one vote, and decide which models it offers and what it charges. The computing itself runs on the co-op's own central infrastructure.

Shared AI tries to answer both at once: run by the people taking part, on the machines they brought with them. That second part is where things get difficult.

The key questions

Can several ordinary computers pool their resources to run useful AI models together, without relying on a central inference server?

And a second one:

What has to be true for that to become useful shared infrastructure, rather than an interesting technical demonstration?

We ended up learning about both.

Loop one: getting it to run

Hypothesis

If a model can be divided into layers, and different devices can run different layers, then a group of ordinary computers should be able to run a model together that would be difficult for any one of them to run alone.

Method

We ran Shared AI with eight participants, each contributing one device. People joined a shared room in their browser and pledged some of their available GPU memory. The system then allocated a section of the model to each participant.

Inference became a pipeline:

host → device A → device B → device C → host

Diagram: a model's layers split into three slices, each held by a different device. Activations pass from the host through devices A, B and C, then back to the host, which picks the next token.
FIG 2 Image credit: FIELD STATION

Each device ran its allocated layers using WebGPU, then passed the activations to the next machine over WebRTC. Eventually they returned to the host, which produced the next token, and the cycle started again. We started with small models to establish that the underlying runtime worked. But those smaller models could have run on a single machine. So we then moved to larger models, where pooling resources starts to matter.

Result

It sort of worked. For an experiment, "sort of" turned out to be quite interesting.

Several devices joined together, received different parts of a model and took part in the same inference process. The largest model we ran, Qwen 27B, produced output. The core idea held up: the model did not need to live on one computer. But not everyone could take part. One participant's browser didn't support WebGPU, so their machine couldn't join the swarm at all.

Getting a model spread across a network was only the beginning. Model preparation took a long time. In one of our diagnostic captures, Llama 3 8B took around 3 minutes 36 seconds to become ready. Qwen 27B took about 12 minutes 21 seconds. And the response time was slow.

That delay matters. A system can be technically distributed and still have enough friction that people give up before it's ready to use.

It's also worth noting that Qwen 27B could actually be run on a single machine, albeit a fairly powerful one, such as a Mac with 32GB unified memory or a good GPU.

Loop two: the slowest machine sets the pace

Hypothesis

If each machine contributes some memory and computation, adding more machines should increase the size of model the group can run, and potentially the useful capacity of the system.

The total resources of the room should matter more than the resources of any individual machine.

Method

We tried the experiment with what a real shared system will always contain: different computers.

Not a carefully matched cluster of GPUs in a data centre, but people's actual machines, with different GPUs, amounts of memory and network connections. Some participants had fast Macs with Apple silicon, some were on Windows, some tried phones.

We instrumented the experiment so we could observe model loading, local execution, network traffic and the movement of activations between peers.

Result

Pooling capacity worked in one sense, but the shape of the network mattered enormously.

A fast computer and a slower one did not average out into a slightly faster collective computer. In a sequential pipeline, every device matters, so each token waits for the slowest machine and the slowest connection.

Timeline diagram: each token is one trip through devices A, B and C and back to the host. The slowest device takes the largest share of every trip, and the next token can't start until the last one finishes. Proportions are illustrative.
FIG 3 Image credit: FIELD STATION

Inference is not a single exchange between devices, which makes a huge difference. Every millisecond spent waiting for small packets of data to pass through the chain adds up.

This obviously changed how we thought about the problem. Rather than:

How much of the model can this machine hold?

Perhaps we should be asking:

Where should this part of the model live so that the complete route produces useful answers quickly?

Memory, compute speed, bandwidth and latency are all part of the same scheduling problem. Dividing a model evenly is probably not enough.

Loop three: people close their browsers

Hypothesis

If computers can join together to run a model, the swarm should behave like a shared resource for as long as enough collective capacity remains available.

Method

For this part, we mostly needed people to behave normally. They did:

Result

This exposed the most important limitation of the experiment.

There was no redundancy.

If one participant held a particular range of model layers, that participant was part of the route. When their browser closed, those layers disappeared with it. That's what happened to our Qwen 27B run: one closed tab, and the whole collective inference process came down.

This wasn't an annoying edge case; it's just real life. We had divided a model between several machines, but we had not made the model resilient to the machines themselves changing.

A collection of people's computers is dynamic by nature. Laptops sleep, Wi-Fi drops, people leave rooms, browsers throttle background tabs, and machines have different capacities. All of this can and will change.

If shared AI depends on everybody staying exactly where they were when inference started, it isn't shared infrastructure yet.

What we learned

The experiment answered our first question positively, with a substantial asterisk.

Yes: several ordinary computers can run an AI model together.

That turned out not to be the most interesting question. The harder one is:

Can the model stay available while the computers underneath it change?

Shared inference is as much a distributed-systems problem as an AI problem: discovery, routing, replication, latency, failure, recovery and redistribution. Projects such as Petals and exo have been working on parts of this for a while. Boring infrastructure stuff, but it decides whether any of this is usable. For shared AI to work, it needs redundancy.

The next loop: RELAY

That's where we're taking the experiment next.

RELAY takes the same underlying idea into a native environment. But it isn't simply Shared AI without the browser, or a rebrand of SwarmLLM. We've actually focused on the harder, more boring, but potentially more useful things:

discovery, routing, latency, redundancy and recovery.

The first RELAY question is:

Can several ordinary computers become a resilient AI system that redistributes itself as the network changes?

Instead of treating the model as something installed on individual machines, RELAY treats it as a shared object that exists across the network. That means peer-to-peer model distribution, replication, re-routing and recovery.

Two-panel diagram. Today: each slice lives on one device, so when device B's tab closes the chain breaks and the answer stops. RELAY: a spare copy of slice B takes over, the route moves to it and the answer carries on.
FIG 4 Image credit: FIELD STATION

Next time somebody closes their laptop mid-answer, we want the model to carry on.

  1. This might not be as true as they claim. An open issue on Darkbloom's repository reports that the end-to-end encryption described in their paper is disabled in production, so the coordinator operator can read prompts and responses. ↩