The Best Local Models Right Now (That I Actually Run)
Forget the cloud — these are the frontier models I run locally every day, what they're actually good at, and the hardware you need.
Everyone talks about "local AI" like it's a hobby. It's not. Running frontier-class models on your own hardware is the biggest cost advantage most businesses are leaving on the table.
Here's the stack I actually run every day — not benchmarks, real work:
The core models
- Qwen 3.5 (27B) — my main agent brain. Handles multi-step reasoning, tool calls, and long context without breaking a sweat. Runs on a single GPU.
- DeepSeek-R1 (14B) — when I need deep chain-of-thought reasoning. The distilled version keeps 90% of the thinking at a fraction of the VRAM.
- Qwen 2.5 VL (7B) — vision. Image QC, reading screenshots, checking generated art. This replaced a paid API for me entirely.
- FLUX.1 (fp8) — image generation that rivals Midjourney for product and marketing work. Runs locally, no per-image fee.
- Kokoro — TTS voiceovers that don't sound robotic. 1GB of RAM, near-zero cost.
What this means for your business
Stop paying per-seat, per-token, per-image. The exact same frontier models the trendy SaaS apps resell to you — you can run them yourself for the cost of electricity. One-time hardware, zero monthly fees, your data never leaves the building.
🎓 Want the full system? The Local Models Course is coming soon — how to pick, run, and profit from frontier AI models on your own hardware, no cloud, no monthly fees. Join the email list and you'll be first to know when it drops.
