46. Local Model Strategy

Choose the right local models (Llama, Qwen, DeepSeek, Mistral) for your hardware and needs.

By Jacques Botte, founder of Toptronic®. Last updated 12 September 2026.

The lesson

Llama 4 Scout (109B active) fits in 24GB VRAM when quantised to 4-bit, making it accessible for high-end consumer hardware.

Llama 4 Maverick (400B) offers deeper reasoning but requires significantly more VRAM or heavy quantisation.

DeepSeek V4 Pro offers 1M context and strong coding/reasoning at dramatically lower cost than Western frontier models.

Qwen 3.7 Max (June 2026) is the new agentic coding flagship with 1M context. Qwen 3.6 models remain excellent for multilingual tasks.

Mistral Large 3 (May 2026) is Apache 2.0 open-weight with 256K context - a strong Western open-weight option.

Match model size to your VRAM: under 12GB use 7B-13B models; 24GB handles 35B-70B quantised; 40GB+ for larger models.

Check yourself

Question 1: When should you prefer Llama 4 Scout over Llama 4 Maverick?
  1. When you need maximum accuracy
  2. When you have limited VRAM (4-bit quantised 17B fits in 24GB) — correct
  3. Only for image tasks
  4. Never

Answer: When you have limited VRAM (4-bit quantised 17B fits in 24GB)

Llama 4 Scout is smaller and fits better on consumer hardware, while Maverick needs more VRAM.

Question 2: What is DeepSeek V4 Pro best suited for locally?
  1. Real-time voice chat
  2. Reasoning, coding, and math with 1M context at lower cost — correct
  3. 8K video editing
  4. Spreadsheet macros only

Answer: Reasoning, coding, and math with 1M context at lower cost

DeepSeek V4 Pro offers strong reasoning and 1M context locally at dramatically lower cost than Western frontier models.

Question 3: Why might you choose Qwen 3.7 Max or 3.6 for a multilingual local application?
  1. It only supports English
  2. Strong multilingual capabilities in a runnable local size — correct
  3. It is the smallest model available
  4. It requires no GPU

Answer: Strong multilingual capabilities in a runnable local size

Qwen models (3.7 Max for coding, 3.6 for multilingual) excel at multilingual tasks while remaining runnable on consumer hardware.

← Previous lesson · All 83 lessons · Next lesson →

The full course — 83 lessons and 249 quiz questions — ships inside the app. Get TPEE to study it offline.