تخطَّ إلى المحتوى

Local LLM Hardware Recommendations

هذا المحتوى غير متوفر بلغتك بعد.

Companion to the Local LLM Hardware Guide. That page covers the high-level decision (local vs API). This page is the per-model hardware matrix you need when buying or renting GPUs.

All tokens-per-second numbers below are batch-size 1, single user, with the listed quantization. Throughput drops 30-50% under multi-user load; bump up a tier for shared deployments.

  • GPUs: RTX 3060 12 GB, RTX 4060 8 GB, RTX 4060 Ti 16 GB
  • Apple Silicon equivalent: M1/M2 base 8-16 GB unified memory
  • Cost: $300-500 used / new
  • Runs well: 7B-13B models at Q4-Q5
  • Don’t try: anything 30B+ — quantization loss kills quality
  • GPUs: RTX 4090 24 GB, RTX 3090 24 GB, RTX A4000 16 GB
  • Apple Silicon equivalent: M3 Pro 18-36 GB, M3 Max 64 GB
  • Cost: $1,000-2,000 used / new
  • Runs well: 30B at Q4, 70B at Q4 (tight, 2-bit weights), Mixtral 8x7B at Q5
  • Sweet spot: 95% of self-host workloads land here
  • GPUs: RTX A6000 48 GB, dual RTX 4090 (NVLink), H100 80 GB, dual A6000
  • Apple Silicon equivalent: M2 Ultra 192 GB, M3 Ultra 192 GB
  • Cost: $4,500-30,000+
  • Runs well: 70B at FP16, 405B at Q4 (multi-GPU), Mixtral 8x22B at Q5, multi-modal vision models
  • When to buy: regulated data, sustained 50M+ token/month workloads, on-prem requirements
ModelParametersRecommended VRAMQuantizationTokens/secTier
Llama 3.1 8B8B6-8 GBQ4_K_M30-60 (4090) / 8-15 (M2 base)1
Llama 3.1 70B70B24-48 GBQ4_K_M15-25 (4090) / 8-12 (M3 Max)2-3
Llama 3.1 405B405B220+ GBQ4_K_M5-10 (8x A100 / M2 Ultra)3
Mistral 7B7B6-8 GBQ4_K_M40-70 (4090)1
Mixtral 8x7B47B (13B active)24-32 GBQ4_K_M30-50 (4090)2
Mixtral 8x22B141B (39B active)80-100 GBQ4_K_M15-25 (A100) / 10-15 (M2 Ultra)3
Qwen 2.5 7B7B6-8 GBQ4_K_M35-65 (4090)1
Qwen 2.5 32B32B20-24 GBQ4_K_M18-28 (4090)2
Qwen 2.5 72B72B40-48 GBQ4_K_M12-20 (A6000)3
DeepSeek-V3671B (37B active)380+ GBQ4_K_M10-15 (8x H100)3+
Phi-4 14B14B10-12 GBQ4_K_M25-45 (4090)1-2

Quantization key: Q4_K_M is the standard 4-bit GGUF format with mixed precision on critical layers. FP16 doubles VRAM needs but improves coherence on edge cases. Drop to Q3_K_S if you must squeeze a tier larger model into limited VRAM, but expect quality loss.

PathBest forThroughput floorThroughput ceiling
CPU only (DDR5)7B emergency fallback1-3 tok/s8 tok/s
Consumer NVIDIA (4060-4090)Single user, mixed sizes8 tok/s70 tok/s
Datacenter NVIDIA (A100, H100)Multi-user batch serving30 tok/s200+ tok/s
Apple Silicon (M3 Max, M2 Ultra)Quiet desk, large models in unified memory6 tok/s30 tok/s
AMD (RX 7900 XTX, MI300X)ROCm-capable workloads10 tok/s80 tok/s

Apple Silicon punches above its weight on 70B+ thanks to unified memory — one M2 Ultra holds models that need a multi-GPU rig on NVIDIA, at the cost of lower peak throughput.

For multi-GPU rigs:

  • PSU: 1000-1600W 80+ Platinum. RTX 4090 peaks near 600W. Dual 4090 needs 1600W minimum with margin.
  • Cooling: front-to-back airflow, minimum 3 intake fans. GPU temps above 80°C throttle inference. Open-frame mining chassis are fine for home rigs.
  • PCIe lanes: prefer x16/x16 over x8/x8 for multi-GPU model sharding. Threadripper or Xeon W gives you the lanes; consumer Intel/AMD desktops don’t.
  • RAM: 64 GB system RAM minimum for 70B+ model loading. Slow DDR4 is fine; the model lives on the GPU.
  • Storage: NVMe SSD with 100+ GB free per model. Llama 3.1 405B at Q4 is ~230 GB on disk.

Single-GPU rigs in a tower case work without modification on a quality 850W PSU.

Before buying, validate the workload on rented hardware. Three reputable options:

ProviderA100 40GBH100 80GBNotes
RunPod$1.19/hr$2.69/hrSpot pricing 30-50% cheaper, hourly billing
Lambda Labs$1.29/hr$2.49/hrReserved instances cheaper, US-only
Vast.ai$0.40-1.00/hr$1.50-2.50/hrMarketplace, variable reliability

Rough monthly costs at 100% utilization: A100 ~$870-940, H100 ~$1,800-1,950. Compare against amortized hardware cost (RTX 4090 used $1,200 / 36 months ≈ $33 + $50 power = $83/mo) before committing.

Rent for: workload sizing, occasional bursts, regulated data with verified-private cloud. Buy for: predictable steady load, air-gapped requirements, multi-year horizon.

Verify on your own hardware before sizing — model releases and inference engine optimizations move the numbers monthly.