📦 What Ollama and open models are
To run an LLM on your machine, you need two things: a program that knows how to load and serve the model, and the model itself. The program is the Ollama; the models are the "open" ones (Qwen, Gemma, Mistral...). This module explains what each one is and how they work together.
📦 What is Ollama
In module 1.2, we separated the LLM (the brain) from the chat interface. But the brain doesn’t run on its own: someone needs to download the model file, load it into memory, and be ready to respond. That “someone” is the Ollama — the program that manages and runs open models on your machine.
New here? Ollama and free software that you install (Mac, Windows, or Linux). Think of it as a "model hub": with it, you download an LLM, chat with the model, and switch models whenever you want—all locally. The technical term for this role is a runtime, in other words, the program that actually makes the model EXECUTE.
Ollama offers two ways to get started, and you can use whichever you prefer: one app with a window (click and chat, like a regular chat) and the terminal (types short commands). Both talk to the same engine under the hood — they’re just two ways to ask for the same thing.
Key concepts
Program that downloads, manages, and runs models locally.
The engine that actually runs the model on your machine.
Two interfaces for the same engine; use whichever you prefer.
Free installation on Mac, Windows, and Linux.
🔓 Open models
Ollama is the program; the open models are what it runs. “Open” here means open weights (open weights): the model file can be downloaded and used by anyone, for free, on their own machine. That's what makes local use possible — without it, you'd always depend on a company's server.
New here? The weights ("weights") are the numbers the model learned during training — its "knowledge," in a file. A model closed (like those from OpenAI) stores those weights on a server, and you access them only through an API. A model of open weights publishes the file: you download it and run it wherever you want. That’s why “open” is the key to the course.
Alibaba's family; the course uses Qwen3 (30B-A3B and 32B versions).
Google's open family; Gemma 3 27B is mentioned.
French models; mentions Mistral Small 3.2 24B.
A strong open family for reasoning, also mentioned.
🌱 Why having multiple models is good
Each family has different strengths (one is better at code, another at text, another is lighter). Since they’re open and free, you can download several, test them, and keep the one that works for you. This freedom to switch is something closed models don’t offer.
Key concepts
With open weights: anyone can download and run it.
The numbers learned during training — the model’s “knowledge.”
Closed stays on the server; open, you download.
Qwen, Gemma, Mistral, DeepSeek — each with its strengths.
⬇️ Download once, run locally
Here's the detail that changes everything about the cost: you downloads the model ONCE. After that, the file lives on your disk and runs locally — no new connection, no new charge. It's the difference between buying a book (pay once, read forever) and renting by the page.
Download (once)
Ollama downloads the model file from the internet. It’s the only step that needs a network connection — and it can be large (some models exceed 15 GB).
Stays on disk
The model becomes a local file on your machine. You can list what you’ve already downloaded with ollama list and delete what you don't use.
Runs offline, for free
From here on, every conversation happens on your machine — without internet or per-use costs, forever.
ollama list
New here? O terminal and it's that text screen where you type commands. Every Ollama command starts with the word ollama followed by what you want (list, run, pull...). You’ll only really use them in Track 2; for now, you just need to recognize the pattern.
Key concepts
Only the first step requires internet access.
The model becomes yours; it takes up space (possibly GBs).
Shows the models already downloaded.
After downloading, every conversation is free.
🗂️ How Ollama serves the model
Here's the piece that connects everything. When Ollama is running, it doesn't just wait for you to open the app: it keeps a local service running in the background, with a address on your own machine. Any program on this computer can send a request to this address and receive the model's response.
New here? One endpoint (or "service address") is like a local doorbell: a fixed place where another program "rings" to ask for something. Ollama's runs on your own machine (at localhost — “this machine here”), so the request never leaves the computer. This is EXACTLY how Hermes will talk to your model in Track 2.
Follow the arrows: the Hermes makes the request to the Ollama service, which loads the model and returns the response—all inside the dashed box (your machine). This is the pipeline you'll connect in module 2.5.
Key concepts
Ollama stays on in the background, ready.
The address where other programs request responses.
"This machine"—the request never leaves the computer.
And this is the endpoint the agent uses to talk to the model.
🔢 Parameters: what does the "B" mean?
You’ll see names like Qwen3 32B or Gemma 3 27B. This "B" is the first thing to understand when choosing a model: it indicates the parameters, in billions. Broadly speaking, more parameters = a more capable model, but also a heavier one.
New here? Parameters are the model’s “internal settings” that it learned — those are the weights from topic 2, now counted. “32B” means 32 billion of them. Since each parameter takes up memory, the “B” is also a clue to how much RAM the model will ask. (RAM = the computer’s working memory; we’ll cover it in module 1.4.)
The bar grows with the “B”: the ~8B one is lightweight and fast, those at 27–32B (the ones in the video) are the balance, and 70B+ models are more capable but require much more memory. "Bigger" isn't always right for YOUR machine.
📊 Read carefully
A bigger "B" isn't automatically better for you: a model that's too large won't fit in your memory and may not even run. The right number is the one that delivers good quality WITHIN what your hardware can handle — covered in module 1.4 and in the practical selection in Track 2.
Key concepts
The internal adjustments learned during training; the “B” counts them in billions.
32 billion parameters, for example.
More parameters require more RAM and run slower.
The best "B" depends on your hardware, not the biggest number.
🆚 Ollama (local) vs. cloud
To round out the vocabulary, it helps to compare them directly. Using a model through Ollama and using a model through the cloud They solve the same problem (generating text), but in very different places and conditions.
✓ Ollama (local)
- ✓Model on your disk; data doesn't leave.
- ✓Free after downloading; runs offline.
- ✓You can freely choose and switch models.
- ✓Limited by your machine’s RAM and chip.
✗ Cloud (API model)
- ✗Data travels to the company's server.
- ✗Charges by usage and requires internet.
- ✗You don’t control the model or changes to its rules.
- ✓In exchange: it’s usually more powerful at difficult tasks.
No ideology: as we saw in 1.1, it's not "Ollama always." The cloud wins at heavy tasks; local wins on privacy, cost, and offline use. Hermes lets you use both—and the three modes in 1.6 are exactly how to switch between them.
Key concepts
Private, free, offline; limited by your hardware.
Powerful, but paid, online, and sends data out.
Both generate text; the location and conditions differ.
Hermes switches between local and cloud (module 1.6).
Optional self-check: What’s the relationship between Ollama and a model like Qwen3?
🎯 Module summary
Next module:
1.4 — Context window and parameters