Skip to content

How Local AI Works on Your Device

When you open ɳClaw for the first time, it looks at your device — the phone you’re holding, the Mac on your desk, the age of your laptop — and automatically picks the right AI model to run locally. This guide explains how that works and what you can control.

Why Did ɳClaw Pick This Model on My Computer?

Section titled “Why Did ɳClaw Pick This Model on My Computer?”

ɳClaw runs a small AI model entirely on your device. That model never leaves your machine. Your conversations stay private.

But not every device can run the same model. A 2015 MacBook can’t run the same AI as an M3 Max MacBook Pro. An iPhone 12 can’t run the same model as an iPhone 16. So ɳClaw picks a model that fits your hardware — one that runs fast and doesn’t overheat.

The choice is automatic. ɳClaw measures your device the first time you open the app and picks the best fit from five tiers.

Each tier targets different hardware. Lighter tiers run on older phones and laptops. Heavier tiers run on newer, more powerful devices.

For older phones with limited RAM or very old netbooks.

  • Model: Qwen 2.5 0.5B
  • Size: ~350 MB
  • Speed target: 15–30 tokens per second
  • Hardware: Android 4GB, iPhone 11/12 in low-power mode, netbooks from 2012–2015

When ɳClaw runs on Tier 0, responses are slower but still usable. The model understands context and can write coherently, just at a measured pace.

For mid-range phones and modest laptops.

  • Model: Llama 3.2 1B
  • Size: ~700 MB
  • Speed target: 20–40 tokens per second
  • Hardware: iPhone 13/14, mid-range Android (6–8GB RAM), 8GB Intel/AMD laptops, older iPad

Tier 1 is the entry point for most devices. Responses feel natural — not instant, but brisk enough to have a conversation.

The most common tier. Most users land here.

  • Model: Llama 3.2 3B
  • Size: ~2 GB
  • Speed target: 25–50 tokens per second (60–100 on Apple Silicon)
  • Hardware: M1/M2 base MacBooks, iPhone 15/16, iPad Air, Snapdragon 8 Gen 2+ Android, modern Windows laptops (16GB RAM)

Tier 2 is where ɳClaw feels truly responsive. On Apple Silicon Macs, it’s noticeably faster. The model runs hot on older hardware but comfortably on modern chips.

For power users with newer workstations or gaming rigs.

  • Model: Llama 3.1 8B
  • Size: ~4.5 GB
  • Speed target: 30–80 tokens per second
  • Hardware: M1/M2 Pro MacBooks, 16–32GB workstations, gaming PCs (8GB+ discrete GPU VRAM)

Tier 3 gives you noticeably smarter responses. The model has more parameters and can reason through longer problems. Battery drain is higher on mobile.

The largest models. You must explicitly choose this.

  • Model: Qwen 2.5 14B or Llama 3.1 70B
  • Size: 8–40 GB
  • Speed target: 15–40 tokens per second
  • Hardware: M2/M3/M4 Max/Ultra, 64GB+ workstations, multi-GPU rigs

Tier 4 is for when you need maximum reasoning. The largest models approach cloud AI quality but run entirely on your machine. Only available if your device has enough disk and memory.

First-Run Benchmark: Why Your Model Might Change

Section titled “First-Run Benchmark: Why Your Model Might Change”

The first time you open ɳClaw, it runs a 60-second benchmark test. This test measures:

  • How many tokens per second your device can generate
  • Peak memory usage
  • Whether the CPU or GPU throttles under load
  • Thermal performance (heat management)

The benchmark process is automatic and runs in the background. You don’t need to do anything.

Why benchmarking matters: A device might look capable on paper (8GB RAM, modern CPU) but throttle when under real load — maybe you’re running many apps, or the device is in a hot environment, or battery is draining fast. The benchmark catches these real-world constraints.

What happens after the benchmark: If your device runs slower than the target for your tier, ɳClaw automatically drops down one tier. For example, if you have an M1 MacBook that should run Tier 2, but the benchmark shows it throttles, ɳClaw picks Tier 1 instead.

If your device runs faster than expected, ɳClaw offers a one-time upgrade prompt: “Your device is faster than expected. Want to try the next tier?” You can accept or decline.

Tier changes happen monthly. ɳClaw re-benchmarks your device every 30 days (or when you restart the app) to stay in sync with real performance.

Mobile Dampers: Power Management on Phones

Section titled “Mobile Dampers: Power Management on Phones”

ɳClaw runs local AI differently on phones than on desktop. Mobile devices are battery-constrained.

Low-power mode: When your phone enters low-power mode (Settings → Battery), ɳClaw automatically drops one tier. This reduces power draw. You can override it if you want faster responses at the cost of battery drain.

Battery under 30%: When your battery drops below 30%, local AI is disabled by default. ɳClaw still works, but it uses a cloud model instead of local inference. If your phone is plugged in, local AI stays active.

Override in Settings: Go to Settings → AI to enable local AI when battery is low. ɳClaw will warn you about battery drain.

These constraints are intentional. We want ɳClaw to be useful across a 12-hour day, not drain your battery in an hour.

Role-Specific Models: When ɳClaw Runs Multiple Models

Section titled “Role-Specific Models: When ɳClaw Runs Multiple Models”

ɳClaw uses local AI for several roles:

  • Chat: Your main conversation — the primary local model
  • Summarizer: Consolidating your memory from earlier conversations
  • Embedder: Creating vector embeddings for memory search (a small, specialized model)
  • Code: (Developer mode) Generating or explaining code

For Tier 0 and 1, all roles use the same model to keep device usage low. Tier 2 and higher can run a separate, smaller embedding model for faster memory search.

You can configure which model runs for each role. Go to Settings → AI → Role Selection. Most users don’t need to change this.

Want to run a different model? You can.

  1. Open Settings → AI → Model Selection
  2. See your current tier and the benchmark score
  3. Choose from presets (Tier 0–4)
  4. Or import a custom GGUF quantization if you have one

Advanced: Users familiar with Ollama can bring their own models. ɳClaw supports any GGUF-format model. Import steps depend on your platform (desktop has an import button; mobile requires a file manager access).

Changing your model takes effect immediately. The app downloads the new model in the background if you pick a different tier. Download size and time are shown before you confirm.

Privacy: Your Data Never Leaves Your Device

Section titled “Privacy: Your Data Never Leaves Your Device”

This is the core promise. Local AI means:

  • Conversations do not go to any server unless you explicitly choose cloud AI (Settings → Privacy → Cloud Model Fallback)
  • Your personal data, memory, and context stay encrypted on your device
  • ɳClaw never trains on your data
  • If you turn off cloud fallback, you’re fully local — zero data leaves your phone or computer

Optional server-side models are a separate setting. You control which features use cloud vs. local.

Can I run a bigger model than my tier recommends?

Section titled “Can I run a bigger model than my tier recommends?”

Yes. Go to Settings → AI → Model Selection and pick a larger tier. If your device can’t handle it, you’ll see slowdowns or battery drain. Subsequent benchmarks may auto-downgrade you back.

ɳClaw re-benchmarks monthly to adapt to changes: a software update that made your OS slower, a new app you installed that consumes RAM, seasonal temperature changes. The benchmark keeps your tier matched to real performance.

Running AI on-device uses more power than cloud AI. Tier 0 uses minimal power; Tier 3+ will drain a phone battery noticeably faster on prolonged use. The mobile dampers (low-power mode, <30% battery) help extend battery life. On desktop, this is not a concern.

Can I disable local AI entirely and use only cloud models?

Section titled “Can I disable local AI entirely and use only cloud models?”

Yes. Go to Settings → Privacy → Inference → Cloud Only. Your conversations then go to your nSelf backend (or ɳClaw’s cloud if you’re using our SaaS). Local models are never invoked.

Llama 3.2 (0.5B, 1B, 3B) and Llama 3.1 (8B) from Meta. Qwen 2.5 (0.5B, 14B) from Alibaba Cloud. All models are permissively licensed. Alternative models can be imported as GGUF files.

ɳClaw detects NVIDIA (CUDA), AMD (HIP), and Intel (oneAPI) GPUs on Linux. It automatically uses GPU acceleration when available. You can see GPU status in Settings → Device Info.