Local LLM Hardware Recommendations
هذا المحتوى غير متوفر بلغتك بعد.
Local LLM Hardware Recommendations
Section titled “Local LLM Hardware Recommendations”Companion to the Local LLM Hardware Guide. That page covers the high-level decision (local vs API). This page is the per-model hardware matrix you need when buying or renting GPUs.
All tokens-per-second numbers below are batch-size 1, single user, with the listed quantization. Throughput drops 30-50% under multi-user load; bump up a tier for shared deployments.
Hardware Tiers
Section titled “Hardware Tiers”Tier 1. Entry — 8 GB VRAM
Section titled “Tier 1. Entry — 8 GB VRAM”- GPUs: RTX 3060 12 GB, RTX 4060 8 GB, RTX 4060 Ti 16 GB
- Apple Silicon equivalent: M1/M2 base 8-16 GB unified memory
- Cost: $300-500 used / new
- Runs well: 7B-13B models at Q4-Q5
- Don’t try: anything 30B+ — quantization loss kills quality
Tier 2. Recommended — 16-24 GB VRAM
Section titled “Tier 2. Recommended — 16-24 GB VRAM”- GPUs: RTX 4090 24 GB, RTX 3090 24 GB, RTX A4000 16 GB
- Apple Silicon equivalent: M3 Pro 18-36 GB, M3 Max 64 GB
- Cost: $1,000-2,000 used / new
- Runs well: 30B at Q4, 70B at Q4 (tight, 2-bit weights), Mixtral 8x7B at Q5
- Sweet spot: 95% of self-host workloads land here
Tier 3. Pro — 48 GB+ VRAM
Section titled “Tier 3. Pro — 48 GB+ VRAM”- GPUs: RTX A6000 48 GB, dual RTX 4090 (NVLink), H100 80 GB, dual A6000
- Apple Silicon equivalent: M2 Ultra 192 GB, M3 Ultra 192 GB
- Cost: $4,500-30,000+
- Runs well: 70B at FP16, 405B at Q4 (multi-GPU), Mixtral 8x22B at Q5, multi-modal vision models
- When to buy: regulated data, sustained 50M+ token/month workloads, on-prem requirements
Per-Model Recommendations
Section titled “Per-Model Recommendations”| Model | Parameters | Recommended VRAM | Quantization | Tokens/sec | Tier |
|---|---|---|---|---|---|
| Llama 3.1 8B | 8B | 6-8 GB | Q4_K_M | 30-60 (4090) / 8-15 (M2 base) | 1 |
| Llama 3.1 70B | 70B | 24-48 GB | Q4_K_M | 15-25 (4090) / 8-12 (M3 Max) | 2-3 |
| Llama 3.1 405B | 405B | 220+ GB | Q4_K_M | 5-10 (8x A100 / M2 Ultra) | 3 |
| Mistral 7B | 7B | 6-8 GB | Q4_K_M | 40-70 (4090) | 1 |
| Mixtral 8x7B | 47B (13B active) | 24-32 GB | Q4_K_M | 30-50 (4090) | 2 |
| Mixtral 8x22B | 141B (39B active) | 80-100 GB | Q4_K_M | 15-25 (A100) / 10-15 (M2 Ultra) | 3 |
| Qwen 2.5 7B | 7B | 6-8 GB | Q4_K_M | 35-65 (4090) | 1 |
| Qwen 2.5 32B | 32B | 20-24 GB | Q4_K_M | 18-28 (4090) | 2 |
| Qwen 2.5 72B | 72B | 40-48 GB | Q4_K_M | 12-20 (A6000) | 3 |
| DeepSeek-V3 | 671B (37B active) | 380+ GB | Q4_K_M | 10-15 (8x H100) | 3+ |
| Phi-4 14B | 14B | 10-12 GB | Q4_K_M | 25-45 (4090) | 1-2 |
Quantization key: Q4_K_M is the standard 4-bit GGUF format with mixed precision on critical layers. FP16 doubles VRAM needs but improves coherence on edge cases. Drop to Q3_K_S if you must squeeze a tier larger model into limited VRAM, but expect quality loss.
CPU vs GPU vs Apple Silicon
Section titled “CPU vs GPU vs Apple Silicon”| Path | Best for | Throughput floor | Throughput ceiling |
|---|---|---|---|
| CPU only (DDR5) | 7B emergency fallback | 1-3 tok/s | 8 tok/s |
| Consumer NVIDIA (4060-4090) | Single user, mixed sizes | 8 tok/s | 70 tok/s |
| Datacenter NVIDIA (A100, H100) | Multi-user batch serving | 30 tok/s | 200+ tok/s |
| Apple Silicon (M3 Max, M2 Ultra) | Quiet desk, large models in unified memory | 6 tok/s | 30 tok/s |
| AMD (RX 7900 XTX, MI300X) | ROCm-capable workloads | 10 tok/s | 80 tok/s |
Apple Silicon punches above its weight on 70B+ thanks to unified memory — one M2 Ultra holds models that need a multi-GPU rig on NVIDIA, at the cost of lower peak throughput.
Cooling, PSU, Chassis
Section titled “Cooling, PSU, Chassis”For multi-GPU rigs:
- PSU: 1000-1600W 80+ Platinum. RTX 4090 peaks near 600W. Dual 4090 needs 1600W minimum with margin.
- Cooling: front-to-back airflow, minimum 3 intake fans. GPU temps above 80°C throttle inference. Open-frame mining chassis are fine for home rigs.
- PCIe lanes: prefer x16/x16 over x8/x8 for multi-GPU model sharding. Threadripper or Xeon W gives you the lanes; consumer Intel/AMD desktops don’t.
- RAM: 64 GB system RAM minimum for 70B+ model loading. Slow DDR4 is fine; the model lives on the GPU.
- Storage: NVMe SSD with 100+ GB free per model. Llama 3.1 405B at Q4 is ~230 GB on disk.
Single-GPU rigs in a tower case work without modification on a quality 850W PSU.
Cloud GPU Rental Fallback
Section titled “Cloud GPU Rental Fallback”Before buying, validate the workload on rented hardware. Three reputable options:
| Provider | A100 40GB | H100 80GB | Notes |
|---|---|---|---|
| RunPod | $1.19/hr | $2.69/hr | Spot pricing 30-50% cheaper, hourly billing |
| Lambda Labs | $1.29/hr | $2.49/hr | Reserved instances cheaper, US-only |
| Vast.ai | $0.40-1.00/hr | $1.50-2.50/hr | Marketplace, variable reliability |
Rough monthly costs at 100% utilization: A100 ~$870-940, H100 ~$1,800-1,950. Compare against amortized hardware cost (RTX 4090 used $1,200 / 36 months ≈ $33 + $50 power = $83/mo) before committing.
Rent for: workload sizing, occasional bursts, regulated data with verified-private cloud. Buy for: predictable steady load, air-gapped requirements, multi-year horizon.
Reference Benchmarks
Section titled “Reference Benchmarks”Verify on your own hardware before sizing — model releases and inference engine optimizations move the numbers monthly.
- vLLM benchmarks: https://docs.vllm.ai/en/latest/performance_benchmark/benchmarks.html
- llama.cpp performance discussions: https://github.com/ggerganov/llama.cpp/discussions (search “performance”)
- Apple MLX benchmarks: https://github.com/ml-explore/mlx-examples
- Anyscale LLM cost vs quality study: https://www.anyscale.com/blog/llama-2-is-about-as-factually-accurate-as-gpt-4-for-summaries-and-is-30x-cheaper
- Ollama model registry with sizes: https://ollama.com/library
See Also
Section titled “See Also”- Local LLM Hardware Guide — when to choose local vs API
- AI cost optimization — Multi-AI Router routing rules
- Plugins / AI — install guide for
plugins-pro/paid/ai - Configuration —
aiplugin settings and provider config