AI Workstation: Maximum Power for 70B+ Models
What this article covers
- Concrete hardware recommendations for 70B and larger models, including dual-GPU setups
- Build suggestions starting at €3000 with specific components and pricing
- Key concepts like NVLink, VRAM pooling, ECC RAM, and PCIe lanes explained clearly
- Decision framework: when a workstation pays for itself versus cloud APIs
- Common pitfalls when building and running an AI workstation
Introduction: Understanding AI Workstations
Anyone trying to run 70B models locally quickly discovers that standard desktop PCs hit a wall. A single RTX 4090 with 24 GB VRAM handles quantized 70B models, but struggles with multiple concurrent users or large agent systems. This is where a proper AI workstation begins: multiple GPUs, substantial RAM, a processor with sufficient PCIe lanes, and a power supply stable enough to support it all.
This article targets developers, researchers, and teams running local AI for production, agent workflows, or multi-model serving. If you’re just starting out, consider reading the AI PC Beginner or AI PC Mid-Range articles first. Foundational concepts are covered in AI Hardware Basics.
Why do you need a workstation?
Picture this: you’re running an agent stack with multiple models simultaneously. A 70B model for reasoning, a smaller one for tool calls, and a vision model for document analysis. On a single consumer PC, this either doesn’t run or crawls. A workstation with two or more GPUs distributes the load, keeps enough VRAM on hand, and enables genuine multi-model serving.
Another scenario: multiple team members accessing local models at the same time. A workstation handles parallel requests without latency skyrocketing. Cloud APIs become expensive fast when data volume grows or sensitive information can’t leave your infrastructure.
A workstation makes sense if you need:
- Local execution of 70B to 120B models at usable speeds
- Multi-user access to local models with low latency
- Agent systems running multiple models concurrently
- Complete data sovereignty, especially with confidential customer information
- Long training runs or fine-tuning your own models
AI Workstation explained
An AI workstation is a high-performance machine designed specifically for running large language models. It features multiple GPUs, substantial VRAM, plenty of system RAM, and a processor with many PCIe lanes. Unlike a standard PC, it executes 70B and larger models smoothly and serves multiple users simultaneously.
Typical specs: dual-GPU setup with 48 GB or more VRAM combined, 128 GB RAM, Threadripper or EPYC processor, and a 1200W or larger power supply. Prices start around €3000 and climb well beyond €10,000.
Who should get a workstation?
A workstation suits users bumping against the limits of consumer hardware. That typically includes:
- Developers and researchers running 70B or larger models locally
- Teams sharing local models and needing low latency
- Agent developers operating multiple models at once
- Enterprises that can’t move sensitive data to the cloud
- Creative studios running image or video generation locally
If you’re only occasionally running a 7B or 13B model, you don’t need a workstation. A beginner or mid-range PC suffices. For guidance on sizing, see Dimensioning Hardware Correctly.
Key workstation terminology
| Term | Meaning |
|---|---|
| Dual-GPU | Two graphics cards in one system providing combined VRAM |
| NVLink | Direct connection between two NVIDIA GPUs for higher throughput |
| VRAM Pooling | Combining VRAM from multiple GPUs into logical shared memory |
| Threadripper | AMD processor with many cores and PCIe lanes, ideal for workstations |
| EPYC | AMD server processor, more cores and lanes than Threadripper |
| ECC RAM | Error-Correcting Code memory that automatically fixes bit errors |
| PCIe Lanes | Data connections between processor and expansion cards, critical with multiple GPUs |
| TDP | Thermal Design Power, maximum heat output of a component |
| PSU | Power Supply Unit for the machine |
| Water Cooling | Liquid cooling, essential for high TDP and multiple GPUs |
| Rack | Server cabinet for very large setups or multiple workstations |
What can a workstation do?
| Scenario | What’s possible | Speed |
|---|---|---|
| 70B quantized (Q4) | Runs on a single 4090 with 24 GB, comfortable on dual-GPU | 15 to 30 tokens/s |
| 70B FP16 | Needs around 140 GB VRAM, only with multiple pro GPUs | 5 to 15 tokens/s |
| 120B quantized | Requires about 70 GB VRAM, dual-GPU or more | 8 to 18 tokens/s |
| Multi-model serving | Multiple models running simultaneously, distributed across GPUs | Depends on load distribution |
| Fine-tuning (LoRA) | Possible with 24 to 48 GB VRAM for 7B to 13B models | Hours to days |
| Image generation | Stable Diffusion and successors run smoothly | Seconds per image |
Exact values depend on quantization, batch size, and framework. For more on quantization, see Quantization. To understand memory bandwidth, read Memory Bandwidth.
GPU configurations
The GPU is the heart of an AI workstation. Here are the most common setups:
Single RTX 4090 (24 GB)
A single RTX 4090 with 24 GB VRAM marks the entry into the workstation space. It handles quantized 70B models and is fast enough for interactive use. For multi-user or agent stacks, it’s insufficient. Expect to pay around €1800 to €2000.
Dual RTX 4090 (48 GB)
Two RTX 4090 cards deliver 48 GB VRAM combined. This comfortably runs 70B models with higher quantization or multiple models in parallel. You’ll need a robust power supply, at least 1200W and ideally 1600W, plus a motherboard with adequate spacing between slots. Without NVLink support (the 4090 lacks it), communication happens over PCIe, which is slightly slower for cross-model workflows.
RTX 6000 Ada (48 GB)
The RTX 6000 Ada is a professional GPU with 48 GB VRAM on a single card. It’s expensive, around €7000 to €9000, but runs very quietly and efficiently. Two of them give 96 GB VRAM, sufficient for 120B models or FP16 inference of 70B. For pure inference, it’s often better than two consumer GPUs.
Mac Studio M2 Ultra (192 GB)
Apple Silicon is its own category. The Mac Studio M2 Ultra with 192 GB Unified Memory can load very large models because the CPU and GPU share the same memory pool. It won’t match NVIDIA GPU speed, but the memory capacity is unbeatable in this price range. Learn more at Unified Memory.
CPU and Motherboard
In a workstation running multiple GPUs, the CPU matters not just for compute power but especially for PCIe lanes. Consumer processors typically offer only 20 to 24 lanes, which gets tight with two GPUs.
AMD Threadripper is the go-to choice for workstations. The current generation offers 48 to 80 PCIe lanes, enough for two to four GPUs at full bandwidth. Models like the Threadripper Pro 7965WX and 7975WX are popular for AI workstations.
AMD EPYC pushes even further but targets servers more than workstations. EPYC processors deliver 96 to 128 PCIe lanes and make sense if you’re running three or four GPUs.
Intel Xeon is an alternative but sees less adoption in AI. The key point: choose a motherboard that supports your desired number of GPUs with adequate slot spacing. Tight slots create heat problems.
RAM and Storage
For 70B models and multi-model serving, you’ll need significantly more RAM than VRAM. Rule of thumb: at least double your VRAM, preferably more.
- 128 GB ECC RAM is the minimum for a serious workstation. ECC corrects memory errors and matters during long inference runs.
- 256 GB RAM is recommended if you’re loading multiple models simultaneously or processing large datasets.
- NVMe RAID speeds up loading large models. Two to four NVMe SSDs in RAID 0 hit several GB/s and noticeably cut load times.
Confirm your motherboard supports ECC. Threadripper and EPYC have it by default; consumer boards often don’t.
Complete Build Suggestions
Here are three concrete builds for different needs and budgets.
Build 1: Dual-4090 Workstation (approx. 4500 EUR)
| Component | Recommendation | Approximate Cost |
|---|---|---|
| GPU 1 | RTX 4090 24 GB | 1800 EUR |
| GPU 2 | RTX 4090 24 GB | 1800 EUR |
| CPU | AMD Threadripper 7960X | 1500 EUR |
| Motherboard | TRX50 board with 2x PCIe 5.0 x16 | 600 EUR |
| RAM | 128 GB DDR5 ECC | 400 EUR |
| SSD | 2 TB NVMe Gen4 | 150 EUR |
| PSU | 1600 Watt 80 Plus Titanium | 300 EUR |
| Case | Large tower with good cooling | 200 EUR |
| Cooling | AIO liquid cooling for CPU | 150 EUR |
| Total | approx. 6900 EUR |
This is the classic professional workstation. 48 GB VRAM handles 70B models with good quantization and multi-model serving. Invest in proper cooling, two 4090s generate serious heat.
Build 2: Single-4090 Workstation (approx. 3000 EUR)
| Component | Recommendation | Approximate Cost |
|---|---|---|
| GPU | RTX 4090 24 GB | 1800 EUR |
| CPU | AMD Ryzen 9 7950X | 600 EUR |
| Motherboard | X670E board | 300 EUR |
| RAM | 64 GB DDR5 | 200 EUR |
| SSD | 2 TB NVMe Gen4 | 150 EUR |
| PSU | 1000 Watt 80 Plus Gold | 150 EUR |
| Case | ATX tower with good airflow | 100 EUR |
| Total | approx. 3300 EUR |
This build offers entry into the workstation class. A single 4090 handles quantized 70B models and individual agent workflows. It’s less suited for multi-user setups.
Build 3: RTX 6000 Ada Professional Workstation (approx. 12000 EUR)
| Component | Recommendation | Approximate Cost |
|---|---|---|
| GPU 1 | RTX 6000 Ada 48 GB | 8000 EUR |
| GPU 2 | RTX 6000 Ada 48 GB | 8000 EUR |
| CPU | Threadripper Pro 7975WX | 3000 EUR |
| Motherboard | WRX90 board | 1000 EUR |
| RAM | 256 GB DDR5 ECC | 800 EUR |
| SSD | 4 TB NVMe Gen5 | 400 EUR |
| PSU | 2000 Watt redundant | 500 EUR |
| Case | Rack or large tower | 400 EUR |
| Total | approx. 22100 EUR |
This build targets 24/7 professional use. 96 GB VRAM accommodates 120B models, FP16 inference for 70B, and genuine multi-model serving. ECC RAM and redundant power supplies ensure stability.
Recommended Hardware at Amazon
Workstation Hardware for Local AI at Amazon
Bei Amazon ansehenAffiliate-Link: Bei einem Kauf erhalten wir möglicherweise eine Provision.
Workstation vs. Cloud
A common question: when does a workstation beat a cloud API? It depends on your usage pattern.
Cloud APIs are cheaper if:
- You call models only occasionally
- You want to test different models without tying up hardware
- Your data can leave your infrastructure
- You don’t want to handle maintenance or hardware management
A workstation pays off if:
- You run large models regularly and API costs climb
- Data sovereignty matters and sensitive data must stay local
- You need low latency for agent systems
- You plan to fine-tune or train models
- Multiple users access it simultaneously
A simple calculation: if you spend 200 EUR per month on cloud APIs, a 4000-EUR workstation breaks even in about 20 months. Add electricity and maintenance, but you keep full control.
Common Workstation Pitfalls
- Undersized power supply: Two 4090s draw over 900 watts combined under load. A 1000-watt unit is tight; 1600 watts is safer.
- Inadequate cooling: Multiple GPUs produce significant heat. Without proper case ventilation or liquid cooling, thermal throttling will occur.
- Insufficient PCIe lanes: Consumer CPUs share lanes, reducing inter-GPU bandwidth. Threadripper or EPYC make sense for dual-GPU setups.
- No ECC RAM: Long inference runs risk silent memory errors leading to unstable results. ECC is standard on professional workstations.
- Poor slot spacing on the motherboard: Two large GPUs need room. Tight slots cause heat buildup and fan throttling.
- Home electrical load: A workstation with a 1600-watt supply can strain a standard outlet. Check your circuit protection, especially under continuous load.
- NVLink not on every GPU: The RTX 4090 lacks NVLink. For VRAM pooling via NVLink, step up to professional GPUs like the RTX 6000 Ada.
- Software compatibility: Not every framework handles multi-GPU equally well. Verify beforehand that your stack can work with distributed VRAM.
Hardware, Costs, and Safety with Workstations
A workstation is an investment that demands careful planning. Beyond purchase price, account for ongoing costs: electricity, maintenance, and possible replacement parts. A dual-4090 workstation draws roughly 700 to 900 watts under load, which adds up with continuous operation.
On the security side, a workstation has an edge over the cloud: your data stays with you. This matters for customer data, internal documents, or research datasets. Still implement regular backups, preferably to a NAS or encrypted cloud service.
Using ECC RAM reduces the risk of silent memory corruption. During long training runs or fine-tuning, that’s a genuine safety gain.
Further Reading and Workstation Resources
- Overview of all buying guides
- Entry-level AI PC to get started
- Mid-range AI PC for moderate workloads
- AI Hardware Fundamentals for theoretical background
- Sizing Hardware Correctly for planning purposes
- Memory Bandwidth as a performance factor
- Unified Memory as an alternative to dedicated VRAM
- Quantization for efficient VRAM usage
- Ollama as a lightweight runtime
FAQ: AI Workstations - Common Questions
What’s the difference between an AI workstation and a regular PC?
An AI workstation has multiple GPUs, significantly more VRAM and RAM, a CPU with many PCIe lanes, and a robust power supply. It’s designed to run large models locally and smoothly, whereas a standard PC lacks the memory and compute power for this.
Is an RTX 4090 enough for 70B models?
Yes. With quantization (Q4 or Q5), a 70B model fits within 24 GB VRAM and runs at interactive speeds. For multi-user scenarios or running multiple models concurrently, you’ll need additional VRAM.
Do I need NVLink for dual-GPU setups?
NVLink accelerates GPU-to-GPU communication but isn’t mandatory. Without it, data moves over PCIe, which is somewhat slower. The RTX 4090 doesn’t support NVLink; only professional GPUs like the RTX 6000 Ada do.
How much RAM does an AI workstation need?
At least 128 GB, ideally 256 GB. A rough rule of thumb: double the total VRAM. If you’re loading multiple models simultaneously or doing fine-tuning, aim for 256 GB.
Is ECC RAM worthwhile for AI?
Yes, for long runs, fine-tuning, or production workloads. ECC automatically corrects memory errors and prevents silent data corruption. On Threadripper and EPYC systems, ECC is standard.
Can I run a workstation in my living room?
Technically possible, but not ideal. Two 4090s under load are loud and generate significant heat. For quiet environments, mini-PCs with Ryzen AI Max or Apple Silicon are better choices.
How much power does a dual-4090 workstation consume?
Under full load, roughly 700 to 900 watts. With peripherals and PSU losses, expect closer to 1000 watts. At idle, consumption is much lower, but sustained full-load operation will noticeably impact your electricity bill.
When does a workstation pay for itself versus cloud APIs?
Typically after 18 to 24 months of regular use. If you spend 200 EUR or more monthly on API calls, a workstation is usually the cheaper option.
Is the Mac Studio M2 Ultra a viable alternative?
For very large models, yes, thanks to 192 GB of Unified Memory. However, throughput is lower than NVIDIA GPUs. For CUDA-optimized frameworks, macOS is less suitable.
Do I need liquid cooling for a workstation?
Not strictly necessary, but recommended with two or more GPUs. Liquid cooling reduces noise and maintains more stable temperatures. Air cooling works but requires a large case with excellent ventilation.
Can I train models on a workstation?
Yes, at least LoRA fine-tuning for 7B to 13B models is feasible with 24 to 48 GB VRAM. Full training of larger models demands significantly more resources and typically requires server-grade hardware.
Which operating system suits an AI workstation?
Linux is the standard for AI workloads because most frameworks and CUDA tools are optimized for it. Windows with WSL2 works too but isn’t the first choice for production.
References and Further Reading
- NVIDIA RTX 4090 Specifications
- NVIDIA RTX 6000 Ada Specifications
- AMD Threadripper Pro Overview
- AMD EPYC Overview
- Ollama Documentation
- Apple Mac Studio
- Workstation Hardware for Local AI on Amazon


