Planning an AI Cluster
What this article covers
- When an AI cluster makes sense
- Hardware topologies
- Networking and storage
- Software options for distributed inference
- Cost and operational considerations
Introduction: Planning an AI cluster
When a single workstation no longer meets your performance needs or you need redundancy, an AI cluster becomes worth considering. A cluster connects multiple computers to work together. This might mean several GPUs sharing a large model, or different workloads distributed across specialized nodes. However, clusters require more planning: networking, storage, synchronization, and software all need to work in concert.
This article covers what to keep in mind when planning an AI cluster.
Key terminology
- Node: A single computer in the cluster
- Master/Worker: Control and execution roles
- Interconnect: High-speed network between nodes
- NVLink/Infinity Fabric: Fast GPU-to-GPU connections
- MPI: Message Passing Interface
- vLLM/TGI: Inference servers
- Ray: Framework for distributed workloads
- Kubernetes: Container orchestration
- RDMA: Remote Direct Memory Access
- Scheduler: Workload distribution
When do you need a cluster?
- Model size exceeds a single GPU
- High volume of parallel requests
- Redundancy and high availability requirements
- Need to scale beyond a single machine
- Separating training and inference workloads
- Team with multiple users
Hardware topologies
Single-node multi-GPU
Multiple GPUs in one machine. Simple, but limited by power supply, cooling, and budget.
Multi-node multi-GPU
Multiple machines networked together. Requires fast interconnects and supporting software.
Hybrid
A powerful master node with specialized workers, for example dedicated to embeddings or vision tasks.
Networking
- 10 GbE minimum
- 25/100 GbE for faster interconnects
- RDMA with InfiniBand ideal
- Redundant connections
- Low latency
Storage
- Shared storage for models and datasets
- NFS, Ceph, or MinIO
- Local NVMe per node for fast model access
- 50-100 Gb network storage for large models
Software options
| Software | Use case |
|---|---|
| vLLM | High-performance inference |
| TGI (HuggingFace) | Text generation inference |
| Ray | Distributed Python workloads |
| Kubernetes | Container orchestration |
| Slurm | Workload manager for HPC |
| Ollama | Simple local inference per node |
Planning steps
- Define your requirements
- Choose a topology
- Set your budget
- Select hardware
- Plan networking
- Choose storage solution
- Define software stack
- Set up security and monitoring
- Build a test environment
- Scale incrementally
Costs
- Multiple servers and GPUs
- Network switches
- Cooling and power infrastructure
- Storage infrastructure
- Maintenance and operational overhead
Tips
- Test with two nodes before scaling up
- Pay attention to network throughput
- Clarify model size and parallelization strategy
- Don’t choose complex software if a single machine suffices
- Set up monitoring and logging from the start
- Keep documentation
Common pitfalls
- Slow networking: Inter-node communication becomes the bottleneck
- Storage bottleneck: Models can’t be distributed quickly enough
- Wrong software: Ollama alone doesn’t scale automatically
- Underestimating power and cooling needs
- Underestimating maintenance overhead
- No redundancy: Single point of failure
Further reading and resources
- BotServ.de Building your own AI server
- BotServ.de AI workstation
- BotServ.de AI mini PC
- BotServ.de Docker Swarm
FAQ: AI clusters
Can I scale Ollama across multiple machines? Not directly. True clustering requires vLLM, Ray, or Kubernetes.
What does a small cluster cost? Anywhere from a few thousand to over 50,000 euros, depending on requirements.
Is 10 GbE enough? For large model transfers, it’s often insufficient.
What is RDMA? A technology for extremely fast network memory access.
Do I need Kubernetes? Only if you’re managing many containers and workloads.
Sources and further reading
- Ray: https://docs.ray.io/
- vLLM: https://docs.vllm.ai/
- Kubernetes: https://kubernetes.io/docs/home/
- Slurm: https://slurm.schedmd.com/
Summary: Planning an AI cluster
An AI cluster connects multiple computers and GPUs to handle larger models or higher volumes of parallel workloads. What matters most is fast networking, shared storage, appropriate software like vLLM, Ray, or Kubernetes, and sufficient budget for power and cooling. A cluster only makes sense when you have clear requirements to justify it. By scaling gradually and setting up monitoring from day one, you’ll avoid costly misconfiguration.


