PTENES
MODULE 1.3

📦 What Ollama and open models are

To run an LLM on your machine, you need two things: a program that knows how to load and serve the model, and the model itself. The program is the Ollama; the models are the "open" ones (Qwen, Gemma, Mistral...). This module explains what each one is and how they work together.

6
Topics
~30
Minutes
Basic
Level
Theory
Type
1

📦 What is Ollama

In module 1.2, we separated the LLM (the brain) from the chat interface. But the brain doesn’t run on its own: someone needs to download the model file, load it into memory, and be ready to respond. That “someone” is the Ollama — the program that manages and runs open models on your machine.

New here? Ollama and free software that you install (Mac, Windows, or Linux). Think of it as a "model hub": with it, you download an LLM, chat with the model, and switch models whenever you want—all locally. The technical term for this role is a runtime, in other words, the program that actually makes the model EXECUTE.

Ollama offers two ways to get started, and you can use whichever you prefer: one app with a window (click and chat, like a regular chat) and the terminal (types short commands). Both talk to the same engine under the hood — they’re just two ways to ask for the same thing.

Ollama app with the model selector open: the 'Find model...' field lists models such as qwen3, gemma, and others to download and use
Video frame: the Ollama app with the dropdown "Find model..." open. Notice the list (qwen3, gemma...)—these are the “open models” in the next topic, ready to download with one click. It’s the friendliest entry point for anyone who doesn’t want to use a terminal.

Key concepts

Ollama

Program that downloads, manages, and runs models locally.

Runtime

The engine that actually runs the model on your machine.

App + terminal

Two interfaces for the same engine; use whichever you prefer.

Free

Free installation on Mac, Windows, and Linux.

2

🔓 Open models

Ollama is the program; the open models are what it runs. “Open” here means open weights (open weights): the model file can be downloaded and used by anyone, for free, on their own machine. That's what makes local use possible — without it, you'd always depend on a company's server.

New here? The weights ("weights") are the numbers the model learned during training — its "knowledge," in a file. A model closed (like those from OpenAI) stores those weights on a server, and you access them only through an API. A model of open weights publishes the file: you download it and run it wherever you want. That’s why “open” is the key to the course.

Qwen

Alibaba's family; the course uses Qwen3 (30B-A3B and 32B versions).

Gemma

Google's open family; Gemma 3 27B is mentioned.

Mistral

French models; mentions Mistral Small 3.2 24B.

DeepSeek

A strong open family for reasoning, also mentioned.

🌱 Why having multiple models is good

Each family has different strengths (one is better at code, another at text, another is lighter). Since they’re open and free, you can download several, test them, and keep the one that works for you. This freedom to switch is something closed models don’t offer.

Key concepts

Open model

With open weights: anyone can download and run it.

Weights

The numbers learned during training — the model’s “knowledge.”

Closed vs. open

Closed stays on the server; open, you download.

Families

Qwen, Gemma, Mistral, DeepSeek — each with its strengths.

3

⬇️ Download once, run locally

Here's the detail that changes everything about the cost: you downloads the model ONCE. After that, the file lives on your disk and runs locally — no new connection, no new charge. It's the difference between buying a book (pay once, read forever) and renting by the page.

1

Download (once)

Ollama downloads the model file from the internet. It’s the only step that needs a network connection — and it can be large (some models exceed 15 GB).

2

Stays on disk

The model becomes a local file on your machine. You can list what you’ve already downloaded with ollama list and delete what you don't use.

3

Runs offline, for free

From here on, every conversation happens on your machine — without internet or per-use costs, forever.

Objective: see the models you’ve already downloaded terminal
ollama list
How to verify: a table appears with NAME, SIZE, and MODIFIED. An empty list means you haven’t downloaded any models yet (we’ll do that in Track 2). This is a real command; run it as is—nothing to replace.

New here? O terminal and it's that text screen where you type commands. Every Ollama command starts with the word ollama followed by what you want (list, run, pull...). You’ll only really use them in Track 2; for now, you just need to recognize the pattern.

Key concepts

One-time download

Only the first step requires internet access.

File on disk

The model becomes yours; it takes up space (possibly GBs).

ollama list

Shows the models already downloaded.

$0 usage

After downloading, every conversation is free.

4

🗂️ How Ollama serves the model

Here's the piece that connects everything. When Ollama is running, it doesn't just wait for you to open the app: it keeps a local service running in the background, with a address on your own machine. Any program on this computer can send a request to this address and receive the model's response.

New here? One endpoint (or "service address") is like a local doorbell: a fixed place where another program "rings" to ask for something. Ollama's runs on your own machine (at localhost — “this machine here”), so the request never leaves the computer. This is EXACTLY how Hermes will talk to your model in Track 2.

your machine (localhost) — nothing leaves here Hermes(the agent) request → Ollama servicelocal endpoint loads → model (LLM)in memory ← answer

Follow the arrows: the Hermes makes the request to the Ollama service, which loads the model and returns the response—all inside the dashed box (your machine). This is the pipeline you'll connect in module 2.5.

Key concepts

Local service

Ollama stays on in the background, ready.

Endpoint

The address where other programs request responses.

localhost

"This machine"—the request never leaves the computer.

Bridge to Hermes

And this is the endpoint the agent uses to talk to the model.

5

🔢 Parameters: what does the "B" mean?

You’ll see names like Qwen3 32B or Gemma 3 27B. This "B" is the first thing to understand when choosing a model: it indicates the parameters, in billions. Broadly speaking, more parameters = a more capable model, but also a heavier one.

New here? Parameters are the model’s “internal settings” that it learned — those are the weights from topic 2, now counted. “32B” means 32 billion of them. Since each parameter takes up memory, the “B” is also a clue to how much RAM the model will ask. (RAM = the computer’s working memory; we’ll cover it in module 1.4.)

more “B” → more capable, heavier ~8Blightweight / fast ~27-32Bbalance (from the video) 70B+very capable, requires a lot of RAM

The bar grows with the “B”: the ~8B one is lightweight and fast, those at 27–32B (the ones in the video) are the balance, and 70B+ models are more capable but require much more memory. "Bigger" isn't always right for YOUR machine.

📊 Read carefully

A bigger "B" isn't automatically better for you: a model that's too large won't fit in your memory and may not even run. The right number is the one that delivers good quality WITHIN what your hardware can handle — covered in module 1.4 and in the practical selection in Track 2.

Key concepts

Parameters

The internal adjustments learned during training; the “B” counts them in billions.

32B

32 billion parameters, for example.

Bigger = heavier

More parameters require more RAM and run slower.

Right for you

The best "B" depends on your hardware, not the biggest number.

6

🆚 Ollama (local) vs. cloud

To round out the vocabulary, it helps to compare them directly. Using a model through Ollama and using a model through the cloud They solve the same problem (generating text), but in very different places and conditions.

✓ Ollama (local)

  • ✓Model on your disk; data doesn't leave.
  • ✓Free after downloading; runs offline.
  • ✓You can freely choose and switch models.
  • ✓Limited by your machine’s RAM and chip.

✗ Cloud (API model)

  • ✗Data travels to the company's server.
  • ✗Charges by usage and requires internet.
  • ✗You don’t control the model or changes to its rules.
  • ✓In exchange: it’s usually more powerful at difficult tasks.

No ideology: as we saw in 1.1, it's not "Ollama always." The cloud wins at heavy tasks; local wins on privacy, cost, and offline use. Hermes lets you use both—and the three modes in 1.6 are exactly how to switch between them.

Key concepts

Local (Ollama)

Private, free, offline; limited by your hardware.

Cloud (API)

Powerful, but paid, online, and sends data out.

Same goal

Both generate text; the location and conditions differ.

Coexistence

Hermes switches between local and cloud (module 1.6).

Optional self-check: What’s the relationship between Ollama and a model like Qwen3?

🎯 Module summary

✓
Ollama — the program (runtime) that downloads, manages, and runs local models; app + terminal.
✓
Open models — with open weights (Qwen, Gemma, Mistral, DeepSeek): you download and run them.
✓
Download once + serve locally — then runs offline and for free; the local endpoint is how Hermes talks to the model.
✓
“B” = billions of parameters — more capable and heavier; the right one is the one that fits your machine.

Next module:

1.4 — Context window and parameters