5. Ollama Local AI

Know what Ollama is and how TPEE should describe it safely.

By Jacques Botte, founder of Toptronic®. Last updated 12 September 2026.

The lesson

Ollama is a separate local model runner. It can run models such as Llama, Qwen, Mistral, and other open-weight models on your own PC.

TPEE does not start Ollama, call its HTTP API, download models, or manage the Ollama service. TPEE only helps you write prompts and notes that you may paste into your own Ollama workflow.

Local AI quality depends heavily on RAM, VRAM, quantisation, and model size. Smaller quantised models are faster and cheaper, but may reason less deeply.

Check yourself

Question 1: What is Ollama in this lesson?
  1. A TPEE internal server
  2. A cloud billing provider
  3. A Rust compiler
  4. A separate local model runner — correct

Answer: A separate local model runner

Ollama runs separately from TPEE and can host local models.

Question 2: What does TPEE do with Ollama?
  1. Downloads models automatically
  2. Calls the Ollama HTTP API
  3. Prepares prompts and notes for manual use — correct
  4. Starts the Ollama service

Answer: Prepares prompts and notes for manual use

TPEE must not start or call Ollama; the user runs it separately if desired.

Question 3: What usually limits larger Ollama models first?
  1. Printer ink
  2. RAM or VRAM capacity — correct
  3. Mouse speed
  4. Screen color

Answer: RAM or VRAM capacity

Local model size and KV cache must fit available memory.

← Previous lesson · All 83 lessons · Next lesson →

The full course — 83 lessons and 249 quiz questions — ships inside the app. Get TPEE to study it offline.