46. Local Model Strategy
Choose the right local models (Llama, Qwen, DeepSeek, Mistral) for your hardware and needs.
By Jacques Botte, founder of Toptronic®. Last updated 12 September 2026.
The lesson
Llama 4 Scout (109B active) fits in 24GB VRAM when quantised to 4-bit, making it accessible for high-end consumer hardware.
Llama 4 Maverick (400B) offers deeper reasoning but requires significantly more VRAM or heavy quantisation.
DeepSeek V4 Pro offers 1M context and strong coding/reasoning at dramatically lower cost than Western frontier models.
Qwen 3.7 Max (June 2026) is the new agentic coding flagship with 1M context. Qwen 3.6 models remain excellent for multilingual tasks.
Mistral Large 3 (May 2026) is Apache 2.0 open-weight with 256K context - a strong Western open-weight option.
Match model size to your VRAM: under 12GB use 7B-13B models; 24GB handles 35B-70B quantised; 40GB+ for larger models.
Check yourself
Question 1: When should you prefer Llama 4 Scout over Llama 4 Maverick?
- When you need maximum accuracy
- When you have limited VRAM (4-bit quantised 17B fits in 24GB) — correct
- Only for image tasks
- Never
Answer: When you have limited VRAM (4-bit quantised 17B fits in 24GB)
Llama 4 Scout is smaller and fits better on consumer hardware, while Maverick needs more VRAM.
Question 2: What is DeepSeek V4 Pro best suited for locally?
- Real-time voice chat
- Reasoning, coding, and math with 1M context at lower cost — correct
- 8K video editing
- Spreadsheet macros only
Answer: Reasoning, coding, and math with 1M context at lower cost
DeepSeek V4 Pro offers strong reasoning and 1M context locally at dramatically lower cost than Western frontier models.
Question 3: Why might you choose Qwen 3.7 Max or 3.6 for a multilingual local application?
- It only supports English
- Strong multilingual capabilities in a runnable local size — correct
- It is the smallest model available
- It requires no GPU
Answer: Strong multilingual capabilities in a runnable local size
Qwen models (3.7 Max for coding, 3.6 for multilingual) excel at multilingual tasks while remaining runnable on consumer hardware.
← Previous lesson · All 83 lessons · Next lesson →
The full course — 83 lessons and 249 quiz questions — ships inside the app. Get TPEE to study it offline.