The best PC for AI and machine learning at home is built around one number: VRAM. The short version: an RTX 5090 with 32GB, a Ryzen 9 9950X and 96GB of DDR5 runs 32B language models and Flux for about $5,000. An RTX 5080 with 16GB covers 7B–14B models. Anything at 70B needs 40GB+ of VRAM, which means an RTX PRO card or two GPUs.
The CPU, RAM and storage matter, but they are the supporting cast. If the model does not fit in VRAM, it runs 10–20 times slower on system RAM or does not load at all.
This guide maps model sizes to VRAM, then lays out three build tiers from $2,500 to $10,000+.
Why VRAM Decides Everything in a Local AI PC
A local AI model has to sit entirely in GPU memory to run at useful speed, so VRAM is a hard ceiling rather than a performance curve.
Here is what happens when you load a model:
- The model’s weights are copied into VRAM, and every token generated reads all of them once
- An RTX 5090 reads its 32GB of GDDR7 at roughly 1.8TB/s; dual-channel DDR5 system RAM manages about 90GB/s
- If a model spills over into system RAM, the whole run drops to the speed of the slowest part
- A 7B model that generates 60 tokens per second in VRAM crawls at 3–5 tokens per second when half of it is in RAM
Quantization makes this workable at home. A model stored at Q4 uses about half a byte per parameter, so a 14B model fits in 10GB instead of 28GB.
You also need room for context. Long chats and coding assistants with 32K+ tokens of context add 2–6GB on top of the model.
In a gaming PC, running out of VRAM costs you frames. In an AI PC, it costs you the model.
That is why this guide starts with the GPU and works outward. Our workstation vs gaming PC comparison covers why a gaming build makes a weak AI machine despite sharing the same GPU.
Model Size to VRAM: What Fits on Each GPU
At Q4 quantization, a 7–8B model fits in 8–12GB, 14B needs 10–12GB, 32B needs about 20GB, and 70B needs 40GB or more.
| Model class | Examples | VRAM at Q4 | VRAM at FP16 | GPU that fits it |
|---|---|---|---|---|
| 7–8B | Llama 3.1 8B, Mistral 7B, Qwen 7B | 8–12GB with context | 16GB | RTX 5070 12GB, RTX 5060 Ti 16GB |
| 12–14B | Qwen 14B, Phi-4 14B | 10–12GB | 28GB | RTX 5070 Ti 16GB, RTX 5080 16GB |
| 27–32B | Qwen 32B, Gemma 27B, DeepSeek-R1 32B distill | About 20GB | 64GB | RTX 5090 32GB |
| 70B | Llama 3.3 70B, Qwen 72B | 40GB+ | 140GB | RTX PRO 5000 48GB, RTX PRO 6000 96GB, or 2× RTX 5090 |
| Image generation | Stable Diffusion XL, Flux.1 dev | 12–16GB+ | 24GB+ for Flux at full precision | RTX 5080 16GB minimum, RTX 5090 ideal |
| Fine-tuning (LoRA/QLoRA) | 7–8B base models | 16–24GB | 48GB+ for full fine-tune | RTX 5090, RTX PRO 5000 |
How to read the table:
- Q4 numbers include 2–4GB of headroom for a 16K context window; longer contexts need more
- FP16 is what you need to run a model unquantized, which matters for research and fine-tuning
Where each GeForce card lands:
- RTX 5070 12GB: 7–8B models comfortably, 14B only at tight quantization
- RTX 5070 Ti and RTX 5080 16GB: 14B models, SDXL and Flux, and 7B fine-tuning with QLoRA
- RTX 5090 32GB: 32B models with room for context, Flux at high resolution, and 8B fine-tuning without compromise
Two RTX 5090s give 64GB across two cards, enough for a 70B model split between them. That is a Threadripper build, because it needs 80 PCIe 5.0 lanes and a 1600W PSU.
Pick the largest model you will actually run, add 4GB for context, and buy the GPU that clears that number. 16GB covers 14B models and image generation. 32GB covers 32B models. 70B is a 48–96GB problem.
GeForce vs RTX PRO for AI at Home
A GeForce RTX 5090 is the best-value AI GPU under $2,500; RTX PRO only wins when you need 48–96GB on one card.
What the two lines share:
- Same Blackwell architecture, same CUDA cores and Tensor cores, same PyTorch, llama.cpp, Ollama and ComfyUI support
- Per-GPU inference speed is close between an RTX 5090 and an RTX PRO 5000, because both read memory at similar rates
What RTX PRO adds:
- RTX PRO 4500: 32GB, the same capacity as a 5090 at lower power and in a two-slot card
- RTX PRO 5000: 48GB, enough for a 70B model at Q4 on one card
- RTX PRO 6000 Blackwell: 96GB at around $8,500+, which runs 70B at Q8 or two 32B models at once
- ECC VRAM, which matters for multi-day training runs where a flipped bit corrupts a checkpoint
- Lower power per card, which is what makes a quad-GPU box possible on one circuit
For most home users the answer is an RTX 5090. Its 32GB covers every model class short of 70B at a quarter of the RTX PRO 6000’s price. Our RTX 5090 vs RTX PRO comparison goes deeper on when the professional card earns its cost.
CPU, RAM and Storage for a Machine Learning PC
A 16-core Ryzen 9 9950X with 64–96GB of DDR5 and 4TB of Gen5 NVMe supports any single-GPU AI build; Threadripper is for multi-GPU.
CPU:
- Inference on the GPU barely touches the CPU, but tokenization, data loading and preprocessing pipelines do
- The Ryzen 9 9950X (16C/32T, 5.7GHz boost) handles data prep, runs CPU-offloaded layers and leaves cores free for a browser and IDE
- Threadripper 9970X (32C) or 9980X (64C) is the right move only when you need 80 PCIe 5.0 lanes for two to four GPUs
RAM:
- Rule of thumb: system RAM should be at least twice your VRAM, so a model can be staged and swapped without touching the SSD
- 64GB (2×32GB) is the floor for an RTX 5080 build; 96GB (2×48GB) is the sweet spot for an RTX 5090
- 128–256GB on Threadripper with ECC RDIMM lets you run a 70B model partly in RAM while a second one sits in VRAM
- Our guide to how much RAM a workstation needs covers the AM5 four-stick caveats
Storage:
- Model files are large: a 70B model at Q4 is about 40GB, a Flux checkpoint is 24GB, and a model library crosses 1TB fast
- A 2TB Gen5 OS drive plus a 2TB Gen5 model drive at up to 14GB/s loads a 32B model in a few seconds
Power and cooling:
- An RTX 5090 draws about 575W under load; pair it with a 1200W 80+ Titanium PSU and a native 12V-2×6 connector
- Two RTX 5090s need 1600W and a case with real front-to-back airflow
Size RAM at twice VRAM, give models their own Gen5 drive, and buy a Titanium PSU with 30% headroom over the GPU’s rated draw. Those three choices separate a reliable AI box from one that crashes mid-run.
Three AI PC Builds: Starter, Enthusiast and Lab
These three builds cover the practical range from a first local-LLM machine to a multi-GPU research box.
- CPU Ryzen 7 9700X (8C/16T)
- GPU RTX 5070 Ti 16GB or RTX 5060 Ti 16GB
- RAM 64GB DDR5-6000 (2×32GB)
- Storage 2TB Gen4 NVMe OS + 2TB Gen4 models
- PSU 850W 80+ Gold, ATX 3.1
- Cooling 360mm liquid cooler
This runs 7B–14B chat and coding models through Ollama or LM Studio, generates SDXL and Flux images, and handles QLoRA fine-tuning of 7B models. 16GB of VRAM is the floor worth buying.
- CPU Ryzen 9 9950X (16C/32T)
- GPU RTX 5090 32GB
- RAM 96GB DDR5-6400 (2×48GB)
- Storage 2TB Gen5 OS + 2TB Gen5 models + 2TB datasets
- PSU 1200W 80+ Titanium, ATX 3.1 / 12V-2×6
- Cooling 360mm liquid cooler, high-airflow case
This is the recommended tier. The RTX 5090 runs 32B reasoning models at 30+ tokens per second, Flux at full resolution and 8B fine-tunes without quantization tricks.
- CPU Threadripper 9970X (32C/64T), sTR5
- GPU 2× RTX 5090 32GB (64GB) or 1× RTX PRO 6000 Blackwell 96GB
- RAM 256GB DDR5 ECC RDIMM, quad-channel
- Storage 2TB Gen5 OS + 4TB Gen5 models + 8TB datasets
- PSU 1600W 80+ Titanium, ATX 3.1
- Cooling 360mm liquid cooler, full-tower case with front-to-back airflow
The Lab tier runs 70B models locally, fine-tunes 13B–30B models, and serves two models at once. Threadripper’s 80 PCIe 5.0 lanes give each GPU a full x16 slot. The RTX PRO 6000 route costs more per gigabyte but fits a 70B model on one card with ECC.
If you are weighing the CPU side of the Lab build, our Threadripper vs Ryzen 9 vs Core Ultra 9 guide explains when the lane count justifies the platform cost.
Looking for an AI Workstation?
If you’d rather skip building your own system, Creator Lite is our compact prebuilt workstation: a 16-core Ryzen 9 9950X, an RTX 5080 16GB for 7B–14B models and Flux, 96GB of DDR5 and 4TB of Gen5 NVMe. It’s hand-assembled, cable-managed and burn-in tested for 48 hours under full load before it ships.
- CPU AMD Ryzen 9 9950X (16C/32T)
- GPU PNY GeForce RTX 5080 16GB
- RAM TEAMGROUP T-Create Expert 96GB DDR5
- Storage 2× Samsung 9100 PRO 2TB Gen5 NVMe (4TB)
- PSU be quiet! Dark Power 14 1200W 80+ Titanium
- Case Lian Li A3-mATX
Creator Lite is made to order and ships nationwide in 10–15 business days, fully insured, with a 3-year parts warranty, lifetime support and a 48-hour burn-in. Pay over time with Affirm, with 0% APR options available.
For RTX 5090, dual-GPU and Threadripper configurations sized to a specific model, see our AI workstations page.
Need a different spec — RTX 5090, 128GB+, Threadripper, ECC? Tell us on the workstation request form and you’ll get a parts plan and quote within 24 hours. Or browse all of our prebuilt workstations.
Frequently asked questions
How much VRAM do I need to run local LLMs?▾
At Q4 quantization, 7–8B models need 8–12GB, 14B models need 10–12GB, 32B models need about 20GB and 70B models need 40GB or more. Add 2–6GB for the context window. A 16GB RTX 5080 covers 14B, a 32GB RTX 5090 covers 32B, and 70B needs RTX PRO or two GPUs.
Is an RTX 5090 good enough for AI and machine learning at home?▾
Yes, it is the best-value AI GPU available. Its 32GB of GDDR7 runs 32B models with room for context, generates Flux images at full resolution and fine-tunes 8B models without quantization. It only falls short at 70B, where 40GB+ of VRAM requires an RTX PRO card or a second 5090.
Do I need a Threadripper for an AI PC?▾
Not for a single GPU. A Ryzen 9 9950X with 24 usable PCIe 5.0 lanes feeds one RTX 5090 at full x16 speed. Threadripper becomes necessary with two or more GPUs, because its 80 lanes give each card a full slot and its quad-channel memory supports 256GB or more of ECC RAM.
How much RAM does a machine learning PC need?▾
Plan on at least twice your VRAM. 64GB is the floor for a 16GB RTX 5080 build, 96GB in two 48GB sticks is the sweet spot for an RTX 5090, and dual-GPU or 70B builds want 128–256GB of ECC RDIMM on Threadripper. Extra RAM lets models stage without touching the SSD.
Can I run Stable Diffusion and Flux on a 16GB GPU?▾
Yes. Stable Diffusion XL runs comfortably in 12–16GB, and Flux.1 dev runs on a 16GB RTX 5080 or 5070 Ti using FP8 weights. Full-precision Flux with multiple LoRAs and high-resolution upscaling wants 24GB or more, which is where the 32GB RTX 5090 removes every limit.
CPUBLD is a custom PC builder based in Jersey City, NJ, building gaming PCs, workstations, and servers shipped nationwide. Every build is personally specced, hand-assembled, and burn-tested before it ships.