Local LLM Hardware Recommendations
Local LLM Hardware Recommendations
Section titled “Local LLM Hardware Recommendations”Companion to the Local LLM Hardware Guide. That page covers the high-level decision (local vs API). This page is the per-model hardware matrix you need when buying or renting GPUs.
All tokens-per-second numbers below are batch-size 1, single user, with the listed quantization. Throughput drops 30-50% under multi-user load; bump up a tier for shared deployments.
Hardware Tiers
Section titled “Hardware Tiers”Tier 1. Entry — 8 GB VRAM
Section titled “Tier 1. Entry — 8 GB VRAM”- GPUs: RTX 3060 12 GB, RTX 4060 8 GB, RTX 4060 Ti 16 GB
- Apple Silicon equivalent: M1/M2 base 8-16 GB unified memory
- Cost: $300-500 used / new
- Runs well: 7B-13B models at Q4-Q5
- Don’t try: anything 30B+ — quantization loss kills quality
Tier 2. Recommended — 16-24 GB VRAM
Section titled “Tier 2. Recommended — 16-24 GB VRAM”- GPUs: RTX 4090 24 GB, RTX 3090 24 GB, RTX A4000 16 GB
- Apple Silicon equivalent: M3 Pro 18-36 GB, M3 Max 64 GB
- Cost: $1,000-2,000 used / new
- Runs well: 30B at Q4, 70B at Q4 (tight, 2-bit weights), Mixtral 8x7B at Q5
- Sweet spot: 95% of self-host workloads land here
Tier 3. Pro — 48 GB+ VRAM
Section titled “Tier 3. Pro — 48 GB+ VRAM”- GPUs: RTX A6000 48 GB, dual RTX 4090 (NVLink), H100 80 GB, dual A6000
- Apple Silicon equivalent: M2 Ultra 192 GB, M3 Ultra 192 GB
- Cost: $4,500-30,000+
- Runs well: 70B at FP16, 405B at Q4 (multi-GPU), Mixtral 8x22B at Q5, multi-modal vision models
- When to buy: regulated data, sustained 50M+ token/month workloads, on-prem requirements
Per-Model Recommendations
Section titled “Per-Model Recommendations”| Model | Parameters | Recommended VRAM | Quantization | Tokens/sec | Tier |
|---|---|---|---|---|---|
| Llama 3.1 8B | 8B | 6-8 GB | Q4_K_M | 30-60 (4090) / 8-15 (M2 base) | 1 |
| Llama 3.1 70B | 70B | 24-48 GB | Q4_K_M | 15-25 (4090) / 8-12 (M3 Max) | 2-3 |
| Llama 3.1 405B | 405B | 220+ GB | Q4_K_M | 5-10 (8x A100 / M2 Ultra) | 3 |
| Mistral 7B | 7B | 6-8 GB | Q4_K_M | 40-70 (4090) | 1 |
| Mixtral 8x7B | 47B (13B active) | 24-32 GB | Q4_K_M | 30-50 (4090) | 2 |
| Mixtral 8x22B | 141B (39B active) | 80-100 GB | Q4_K_M | 15-25 (A100) / 10-15 (M2 Ultra) | 3 |
| Qwen 2.5 7B | 7B | 6-8 GB | Q4_K_M | 35-65 (4090) | 1 |
| Qwen 2.5 32B | 32B | 20-24 GB | Q4_K_M | 18-28 (4090) | 2 |
| Qwen 2.5 72B | 72B | 40-48 GB | Q4_K_M | 12-20 (A6000) | 3 |
| DeepSeek-V3 | 671B (37B active) | 380+ GB | Q4_K_M | 10-15 (8x H100) | 3+ |
| Phi-4 14B | 14B | 10-12 GB | Q4_K_M | 25-45 (4090) | 1-2 |
Quantization key: Q4_K_M is the standard 4-bit GGUF format with mixed precision on critical layers. FP16 doubles VRAM needs but improves coherence on edge cases. Drop to Q3_K_S if you must squeeze a tier larger model into limited VRAM, but expect quality loss.
CPU vs GPU vs Apple Silicon
Section titled “CPU vs GPU vs Apple Silicon”| Path | Best for | Throughput floor | Throughput ceiling |
|---|---|---|---|
| CPU only (DDR5) | 7B emergency fallback | 1-3 tok/s | 8 tok/s |
| Consumer NVIDIA (4060-4090) | Single user, mixed sizes | 8 tok/s | 70 tok/s |
| Datacenter NVIDIA (A100, H100) | Multi-user batch serving | 30 tok/s | 200+ tok/s |
| Apple Silicon (M3 Max, M2 Ultra) | Quiet desk, large models in unified memory | 6 tok/s | 30 tok/s |
| AMD (RX 7900 XTX, MI300X) | ROCm-capable workloads | 10 tok/s | 80 tok/s |
Apple Silicon punches above its weight on 70B+ thanks to unified memory — one M2 Ultra holds models that need a multi-GPU rig on NVIDIA, at the cost of lower peak throughput.
Cooling, PSU, Chassis
Section titled “Cooling, PSU, Chassis”For multi-GPU rigs:
- PSU: 1000-1600W 80+ Platinum. RTX 4090 peaks near 600W. Dual 4090 needs 1600W minimum with margin.
- Cooling: front-to-back airflow, minimum 3 intake fans. GPU temps above 80°C throttle inference. Open-frame mining chassis are fine for home rigs.
- PCIe lanes: prefer x16/x16 over x8/x8 for multi-GPU model sharding. Threadripper or Xeon W gives you the lanes; consumer Intel/AMD desktops don’t.
- RAM: 64 GB system RAM minimum for 70B+ model loading. Slow DDR4 is fine; the model lives on the GPU.
- Storage: NVMe SSD with 100+ GB free per model. Llama 3.1 405B at Q4 is ~230 GB on disk.
Single-GPU rigs in a tower case work without modification on a quality 850W PSU.
Cloud GPU Rental Fallback
Section titled “Cloud GPU Rental Fallback”Before buying, validate the workload on rented hardware. Three reputable options:
| Provider | A100 40GB | H100 80GB | Notes |
|---|---|---|---|
| RunPod | $1.19/hr | $2.69/hr | Spot pricing 30-50% cheaper, hourly billing |
| Lambda Labs | $1.29/hr | $2.49/hr | Reserved instances cheaper, US-only |
| Vast.ai | $0.40-1.00/hr | $1.50-2.50/hr | Marketplace, variable reliability |
Rough monthly costs at 100% utilization: A100 ~$870-940, H100 ~$1,800-1,950. Compare against amortized hardware cost (RTX 4090 used $1,200 / 36 months ≈ $33 + $50 power = $83/mo) before committing.
Rent for: workload sizing, occasional bursts, regulated data with verified-private cloud. Buy for: predictable steady load, air-gapped requirements, multi-year horizon.
Reference Benchmarks
Section titled “Reference Benchmarks”Verify on your own hardware before sizing — model releases and inference engine optimizations move the numbers monthly.
- vLLM benchmarks: https://docs.vllm.ai/en/latest/performance_benchmark/benchmarks.html
- llama.cpp performance discussions: https://github.com/ggerganov/llama.cpp/discussions (search “performance”)
- Apple MLX benchmarks: https://github.com/ml-explore/mlx-examples
- Anyscale LLM cost vs quality study: https://www.anyscale.com/blog/llama-2-is-about-as-factually-accurate-as-gpt-4-for-summaries-and-is-30x-cheaper
- Ollama model registry with sizes: https://ollama.com/library
See Also
Section titled “See Also”- Local LLM Hardware Guide — when to choose local vs API
- AI cost optimization — Multi-AI Router routing rules
- Plugins / AI — install guide for
plugins-pro/paid/ai - Configuration —
aiplugin settings and provider config