Skip to content
BotServBotServ
AI ClusterHomelabAI HardwareClusterDistributed AI

AI Clusters for Local AI

Connect multiple computers for AI inference and training. Networking, distributed models, failover, and homelab clusters.

S

schutzgeist

4 min read
AI Clusters for Local AI

AI Clusters for Local AI

What this article covers

  • When a cluster makes sense for AI workloads.
  • The difference between single-node and multi-node inference.
  • Network, storage, and compute requirements.
  • Tools like vLLM, Ray, llama.cpp RPC, and RunPod-like solutions.
  • Tips for building one in a homelab.

Introduction: AI clusters for local AI

A single powerful PC handles most local AI tasks. But once you need to run larger models, ensure higher availability, or distribute inference across hardware, you’ll eventually reach the point where combining multiple machines into a cluster makes sense. An AI cluster can distribute load, split models across multiple GPUs, or pool smaller GPUs together to work as one resource.

This article surveys AI clusters in homelabs and covers what you need to know before building one.

Key terminology

  • Cluster: Multiple machines working together.
  • Node: A single machine in the cluster.
  • Single-node inference: The model runs on one machine.
  • Multi-node inference: The model is distributed across multiple machines.
  • Tensor Parallelism: Splitting a model across multiple GPUs.
  • Pipeline Parallelism: Distributing layers of a model.
  • RPC: Remote Procedure Call; calling functions on distant machines.
  • Load Balancing: Distributing requests across multiple nodes.

When does an AI cluster make sense?

  • A single system cannot hold your target model.
  • You want to combine multiple smaller GPUs.
  • High availability is critical.
  • Different workloads need to run independently.
  • You plan experiments with distributed training or inference.
  • Multiple people need simultaneous access to models.

Hardware basics

Machine per node

Each node needs:

  • A CPU with AVX2 support.
  • RAM equal to at least as much as GPU VRAM.
  • Fast network connectivity, ideally 10 Gbit/s or higher.
  • A GPU with sufficient VRAM.
  • PCIe lanes for GPU and NVMe.

Networking

  • 1 Gbit/s works for small clusters.
  • 10 Gbit/s or higher recommended for larger models.
  • Latency matters more than raw bandwidth.
  • A switch with enough ports and Jumbo Frames is optional.

Storage

  • Shared storage simplifies model distribution.
  • NFS, Ceph, GlusterFS, or plain rsync all work.
  • SSDs are recommended; NVMe for large models.

Types of clusters

Inference clusters

Multiple nodes serve requests. Load distribution happens via reverse proxy or API gateway. Each model can run on its own node.

Model-distribution clusters

A single large model is split across multiple GPUs or nodes. Examples:

  • Tensor Parallelism: Weights are divided.
  • Pipeline Parallelism: Model layers are distributed.

This requires specialized software like vLLM, TensorRT-LLM, or Ray Serve.

Mixed clusters

Some nodes handle inference, others run databases, RAG, or monitoring. This separation improves stability and scalability.

Tools for AI clusters

ToolPurpose
vLLMHigh-performance inference with multi-GPU support.
RayDistributed Python applications and serving.
llama.cpp RPCDistributed inference across the network.
TGIHugging Face Text Generation Inference.
OllamaSimple local inference, primarily single-node.
KubernetesOrchestration for container workloads.
Docker SwarmLighter-weight container orchestration.

llama.cpp RPC

llama.cpp supports RPC to use GPUs on other machines:

./rpc-server -p 50052 -m cuda

On the main machine:

./main -m modell.gguf --rpc 192.168.1.20:50052

This distributes computation across multiple machines. Latency and bandwidth play a large role.

vLLM for multi-GPU

vLLM can distribute models across multiple GPUs in a single machine or across multiple nodes:

python -m vllm.entrypoints.openai.api_server \
  --model meta-llama/Llama-3.1-70B \
  --tensor-parallel-size 2 \
  --pipeline-parallel-size 2

Here, two GPUs share the weights, while two additional nodes execute parts of the pipeline.

Network topologies

  • Star: A central node coordinates everything, simple but a single point of failure.
  • Mesh: Every node communicates with every other node, more robust.
  • Ring: Data flows through a ring, good for specific parallelization strategies.

For homelabs, a star or mesh setup is usually simple enough.

Operating systems and orchestration

  • Proxmox VE: Easy virtualization and containers across multiple nodes.
  • Kubernetes: Powerful but complex.
  • Docker Swarm: Simpler entry into cluster orchestration.
  • Plain Linux with SSH: Minimal overhead, but requires a lot of manual work.

Power, cooling, and space

  • Multiple machines consume significant power.
  • Cooling demands and room temperature increase.
  • Multiple power supplies and UPS systems make sense.
  • Cable management and dust control matter.

Cost

A cluster built from used servers or workstations can be affordable. New hardware with current GPUs is expensive. Often a mixed approach works best: one powerful main machine and cheaper nodes for specialized tasks.

Common pitfalls

  • Network connection too slow: Inference becomes a waiting game.
  • Mismatched GPUs: Load balancing becomes difficult.
  • No shared storage solution: Models must be copied multiple times.
  • Insufficient RAM: Each node must hold the model or its portion.
  • Complex software stack: Debugging becomes tedious.
  • Inadequate power supply: Multiple GPUs draw a lot of current.
  • Cooling underestimated: Server rooms get very hot.

Further reading and resources

FAQ: AI clusters

Does a homelab really need a cluster? No, for most use cases a single powerful machine is enough.

Can I run Ollama on a cluster? Ollama is primarily single-node. For distributed inference, you need tools like vLLM or llama.cpp RPC.

Is 1 Gbit/s network enough? For small models, yes; for large distributed models, probably not.

Can I mix different GPUs? Possible, but not always straightforward. Homogeneous GPUs are much easier to manage.

Is used server hardware worth it? Often yes, especially for CPU workloads and plenty of RAM.

Sources and further reading

Summary: AI clusters for local AI

An AI cluster unlocks larger models, higher availability, and isolated workloads. For homelabs it is not always necessary, but it is a logical next step when a single machine hits its limits. Key priorities are fast networking, sufficient RAM, shared storage, and the right software like vLLM, Ray, or llama.cpp RPC. If you keep power, cooling, and latency in mind, you can build a capable distributed AI system even at a small scale.

Back to Blog
Share:

Related Posts