/00 — boot sequence

Hello.

Article

Mesh LLM: Distributed AI Computing Across Pooled GPUs

July 12, 2026•7 min read
mesh-llm distributed-computing gpu-pooling ai-infrastructure iroh open-source

When people picture running a large language model, they picture a data center. Racks of GPUs that belong to someone else, a metered API, and a bill that grows every month you succeed. You send your prompts off to a black box and hope the price, the model, and the privacy policy all stay the way they were when you signed up.

For a lot of teams that is a bad trade. You give up control over when models change, where your data goes, and what hardware runs your workloads. And as usage grows, so does the bill, with no lever to pull except paying more.

Mesh LLM offers a different shape. It pools the GPUs and memory you already have, across as many machines as you want to add, and exposes the whole thing as one OpenAI-compatible API. Start one node. Add more later. Let the mesh decide whether a model runs on the box in front of you, routes to a peer, or splits across several machines.

The Problem: AI Is Expensive and Centralized

The most popular AI models today are monoliths running in hyperscale data centers. Most developers reach them through an API key and a monthly bill from a large provider. That convenience comes with tradeoffs:

  • No control over model updates. The provider can swap models, deprecate endpoints, or change behavior with no warning.
  • No data privacy. Your prompts and results pass through a third party's infrastructure.
  • No cost leverage. As usage grows, the bill grows linearly with no path to reduce per-token costs by running your own hardware.

Plenty of businesses and individual developers have GPUs sitting in offices, in closets, under desks. What they have been missing is a way to make those machines act as one.

How Mesh LLM Works

Mesh LLM distributes model compute across a mesh of iroh endpoints. A request can be served three ways:

  1. Run it locally, on the machine's own GPU, for latency-sensitive or private workloads.
  2. Route it to a peer that already has the model loaded in memory, avoiding redundant loading.
  3. Split a model too big for any single box across several machines as a pipeline, so layers 0 to 15 run on node A and layers 16 to 31 on node B.

The architecture is pluggable. Plugins declare what they provide in a manifest. The runtime starts them, routes calls, and exposes their capabilities over MCP, HTTP, inference, and mesh events. The catalog ships with over 40 models, from half-a-billion-parameter models that fit on a laptop to 235 billion parameter mixture-of-experts giants.

The Skippy Split Mode

For the largest models, Mesh LLM introduces split mode, internally called Skippy. A model gets partitioned by layer ranges into stages: layers 0 to 15 on one node, 16 to 31 on the next, and so on down the pipeline. Activations flow from one stage to the next, so several modest machines can run a model none of them could hold alone. The OpenAI client never sees any of this. It still just talks to localhost.

This is the key architectural insight: mesh LLM inference turns a hardware problem (not enough VRAM) into a networking problem (fast enough inter-node latency). For teams with multiple mid-range GPUs, this unlocks models that would otherwise require a data center GPU cluster.

Built on Iroh: Peer-to-Peer Networking

The underlying transport layer is iroh, an open-source dial-any-device networking library. Every node in Mesh LLM boots an iroh endpoint. That endpoint is the node's identity, a public key, and its only network surface. There is no central server.

Iroh handles:

  • UDP hole punching for direct connections through NATs and firewalls
  • Relay fallback when direct connections are impossible
  • End-to-end encryption via QUIC, with no need for TLS certificates
  • Public key addressing, so nodes find each other by identity, not IP address

The whole protocol rides on QUIC's ALPN negotiation with three separate protocols: main mesh (gossip, routing, HTTP tunnels, plugin channels), owner control plane (config sync, ownership attestation), and latency-sensitive activation transport for split models.

On a single QUIC connection, everything is a bidirectional stream tagged with a single leading byte that says what kind of stream it is. One connection carries gossip, inference, route queries, and peer lifecycle events, all demultiplexed by that first byte.

Developer Experience

Users install an 18 MB lightweight binary and either join the public mesh or configure private deployments. The system presents itself as an OpenAI-compatible API running on localhost, so any existing OpenAI client library works without modification.

For teams managing their own inference, this means:

  • No per-token API costs beyond electricity
  • Full control over which models run and when they update
  • Data never leaves your network (in private mesh mode)
  • Incremental scaling: add nodes as demand grows

A mobile app using iroh's Swift SDK is in development, and the project plans to support the Agent Communication Protocol (ACP) for multi-agent coordination on the mesh.

Comparing Mesh LLM to Alternatives

FeatureMesh LLMOpenAI APIOllamavLLM
Distributed across machinesYesN/ANoYes
Self-hostedYesNoYesYes
OpenAI-compatible APIYesNativeYesYes
Split large models across GPUsYes (Skippy)N/ANoVia tensor parallelism
Peer-to-peer meshYesNoNoNo
No central serverYesNoYesNo
Cost modelYour hardwarePer-tokenYour hardwareYour hardware

When to Use Mesh LLM

Mesh LLM shines in three scenarios:

1. GPU hoarding. Your team has several machines with consumer GPUs (RTX 4090s, for example) that sit idle most of the time. Mesh LLM pools them into a single inference endpoint.

2. Privacy-sensitive workloads. Healthcare, finance, or legal use cases where sending data to a third-party API is not an option. Run the mesh entirely on-premise.

3. Edge deployments. Offices, labs, or remote locations with unreliable internet. The mesh works offline because peers discover each other directly.

Frequently Asked Questions

Do I need high-speed networking between nodes? Split mode benefits from low latency between nodes, ideally within the same LAN. For local-only routing, standard office networking is sufficient.

How does Mesh LLM handle node failures? In-flight requests to a failed node time out. The mesh rebalances future requests to remaining peers. Persistent state is not stored on mesh nodes.

Can I use Mesh LLM with existing OpenAI code? Yes. The API is wire-compatible with OpenAI's chat completions and embeddings endpoints. Point your client at localhost:11434 or any mesh node IP.

Is Mesh LLM free and open source? Mesh LLM is open source. The iroh library it builds on is also open source, with extended support contracts available for production deployments.

What GPUs are supported? Any GPU supported by llama.cpp, including NVIDIA, AMD, and Apple Silicon. Models must be in GGUF format.

Key Takeaways

  • Mesh LLM pools GPUs across machines into a single OpenAI-compatible API for distributed AI inference.
  • The Skippy split mode lets several modest machines run models too large for any single GPU.
  • Built on iroh, an open-source peer-to-peer networking library using QUIC, hole punching, and public key addressing.
  • Supports over 40 models from 0.5B to 235B parameters with local, peer-routed, and split execution modes.
  • Complete control over data privacy, model selection, and infrastructure costs compared to cloud APIs.

Conclusion

Mesh LLM represents a compelling shift in how developers can think about AI infrastructure. Instead of funneling every prompt through a centralized API, it turns the hardware you already own into a distributed inference mesh. The combination of iroh's peer-to-peer networking, the Skippy split architecture, and OpenAI API compatibility creates a practical path for teams that want more control and lower costs.

For developers tired of rising API bills and opaque model changes, Mesh LLM offers a genuine alternative: use your own GPUs, on your own terms, without sacrificing compatibility.


Source: Iroh Blog - Mesh LLM: distributed AI computing on iroh, Iroh website

Automated Transmission

This entry was synthesized and populated dynamically using native API integrations.

Resources & Links