PTENES
TRAIL 2

🛠️ Hands-on

Enough theory — time to install. You’ll install Ollama on your computer, choose a model that fits your hardware, download it and chat with it, prepare the agent model (Qwen 3 Coder with 64k context), and connect everything to the Hermes Agent — desktop, terminal, or Telegram. Each module includes commands ready to copy and paste.

1 · installOllama 2 · downloada model 3 · create 64kqwen-coder-64k 4 · Hermesconnected local model 100% local $0 · offline

The four practical steps in this track, from left to right: install Ollama, download a model, create the 64k version, and plug everything into the Hermes Agent — arriving at an agent running 100% on your machine.

6
Modules
36
Topics
~3h
Duration
Practical
Level
Learning path progress0%
0 of 36 topics

Track map

Detailed content

2.1~30 min

⬇️ Install Ollama

The first practical step: get Ollama from the website, install it via the terminal (Mac/Linux) or the installer (Windows), and check that everything worked — with commands ready to paste.

0 of 6 · 0%
What it is:

You start on the official ollama.com site, where you’ll find the Download button and the installation command.

Why learn:

Getting it from the official source avoids fake versions and ensures you get the right installer for your system.

Key concepts:

Official source, Download, system detection.

What it is:

On Mac/Linux, a single command (curl ... | sh) downloads and installs Ollama all at once.

Why learn:

It’s the fastest way and the one you’ll actually paste in—without clicking through screens.

Key concepts:

curl, pipe to sh, installation script.

What it is:

On Windows (and Mac too), you can download the installer from the website and click to install.

Why learn:

If you don’t like the terminal, this is the command-line-free path.

Key concepts:

Graphical installer, tray app.

What it is:

Running ollama --version and ollama list confirms it’s installed and responding.

Why learn:

Checking before moving on prevents wasting time later on a "command not found" error.

Key concepts:

Verification, --version, empty list.

What it is:

Ollama has a windowed app and terminal commands—they do the same thing.

Why learn:

Knowing both exist lets you choose whichever feels more comfortable.

Key concepts:

GUI vs CLI, same engine underneath.

What it is:

The most common issue is that the terminal can't "see" the newly installed ollama—just reopen the terminal.

Why learn:

Resolves 90% of installation scares without needing to reinstall anything.

Key concepts:

PATH, reopen the terminal, new session.

View Full
2.2~30 min

🎚️ Choose the right model for your hardware

Before downloading gigabytes, find out what your machine can handle: check the hardware, ask Hermes for a recommendation, understand the effort levels, and learn the rule of leaving RAM headroom.

0 of 6 · 0%
What it is:

Check how much memory (RAM) and which chip you have — on a Mac, in “About This Mac.”

Why learn:

RAM determines which model fits; without that, you're downloading blindly.

Key concepts:

RAM, chip/GPU, “About This Mac”.

What it is:

You show Hermes your hardware, and it suggests models that will fit (e.g., M4 Max 36GB → Qwen 32B / 30B-A3B).

Why learn:

Takes the guesswork out of the choice; the agent itself points you in the right direction.

Key concepts:

Recommendations by hardware, options by size.

What it is:

In the Hermes selector, you adjust the effort (Minimal, Low, Medium, High, Max) and toggle Thinking/Fast.

Why learn:

This is how you balance quality and speed without switching models.

Key concepts:

Effort, Thinking, Fast.

What it is:

The model needs to fit in RAM with room for the rest of the system; if it’s too full, the machine freezes.

Why learn:

Avoids the classic mistake of choosing the largest model and making the machine choke.

Key concepts:

Headroom, RAM buffer, download→test→delete.

What it is:

The q4_K_M suffix indicates a quantized (compressed) model that takes up less memory.

Why learn:

Understanding quantization lets you fit larger models in the same RAM.

Key concepts:

Quantization, q4_K_M, size vs. accuracy.

What it is:

A concrete recommendation to get started: qwen3:30b-a3b-q4_K_M, fast and well-balanced.

Why learn:

Starting with a tested model helps you avoid choice paralysis.

Key concepts:

First model, qwen3:30b-a3b-q4_K_M.

View Full
2.3~30 min

💬 Download and chat with your 1st model

The “wow” moment: download the model, run it in the terminal or app, understand “thinking,” and manage what’s on disk — all running offline on your machine.

0 of 6 · 0%
What it is:

The ollama pull command downloads the chosen model (about 18 GB) to your drive, with a progress bar.

Why learn:

And it’s the only time you need the internet; after that, it runs offline.

Key concepts:

pull, one-time download, progress bar.

What it is:

ollama run opens a chat in the terminal itself; you type, it responds, and you exit with /bye.

Why learn:

It’s the most direct way to prove that the local model works.

Key concepts:

run, prompt in the terminal, /bye.

What it is:

The Ollama app has a chat window; choose a model from the list and chat as you would in a regular app.

Why learn:

It’s the comfortable path for those who prefer the mouse to the keyboard.

Key concepts:

Chat app, model selector.

What it is:

Reasoning models "think" before responding; the app shows something like "Thought for 6.2 seconds".

Why learn:

Explains why the response takes a little while — and why it’s usually better.

Key concepts:

Thinking, reasoning, “thinking” time.

What it is:

The first time you ask a question, the model loads into memory and takes longer; after that, it gets faster.

Why learn:

Knowing this helps you avoid thinking it’s “frozen” right away.

Key concepts:

Cold start, loading into RAM, warm-up.

What it is:

ollama list shows what you downloaded, ollama ps shows what’s running, and ollama rm deletes a model.

Why learn:

You’ll try several; deleting the ones you don’t use frees up disk space.

Key concepts:

list, ps, rm, disk cleanup.

View Full
2.4~30 min

🪟 The agent model: Qwen 3 Coder 64k

The agent requires 64k of context. You’ll understand why, learn about Qwen 3 Coder, write a Modelfile that raises num_ctx, and create the 64k derived model with one command.

0 of 6 · 0%
What it is:

Hermes Agent needs a model with 64,000 context tokens; the one from module 2.3 doesn’t always have that.

Why learn:

It’s why you prepare a specific model before connecting the agent.

Key concepts:

64k requirement, agent model.

What it is:

An open model from the Qwen family designed for coding and tool use—ideal for an agent.

Why learn:

Knowing why it’s chosen helps you switch later with confidence.

Key concepts:

Qwen 3 Coder, a model for agents and tools.

What it is:

A Modelfile is a short recipe: it starts with a base model and sets num_ctx to 65536 (64k).

Why learn:

This is how you “make” the 64k model from one you already have.

Key concepts:

Modelfile, FROM, PARAMETER num_ctx 65536.

What it is:

ollama create qwen3-coder-64k -f Modelfile creates a new model with the context set to 64k.

Why learn:

This derived model is what you’ll point Hermes to in module 2.5.

Key concepts:

create, derived model, -f Modelfile.

What it is:

ollama show qwen3-coder-64k and ollama list confirm that the model exists and has 64k.

Why learn:

Checking before plugging it into the agent prevents context errors down the line.

Key concepts:

show, list, check num_ctx.

What it is:

A larger context uses more RAM; 64k uses more than the same model's default context.

Why learn:

Turn the headroom rule back on: you need spare capacity to run the agent.

Key concepts:

Context memory cost, headroom.

View Full
2.5~30 min

🔌 Connect the local model to the Hermes Agent

The final connection: install/update Hermes (open source, MIT, Nous Research), select the local 64k model, diagnose and test the connection while running 100% offline.

0 of 6 · 0%
What it is:

Hermes Agent is the open-source “AI OS” from Nous Research, under the MIT license.

Why learn:

Open source + MIT = you can run, audit, and adapt it without restrictions.

Key concepts:

Open source, MIT license, Nous Research.

What it is:

hermes update brings Hermes up to the latest version; when it finishes, "HERMES IS READY" appears.

Why learn:

Starting with the latest version helps you avoid bugs that have already been fixed.

Key concepts:

hermes update, "HERMES IS READY".

What it is:

In the Hermes selector, you choose qwen3-coder-64k; the active model appears in the lower-right corner.

Why learn:

It’s the step that makes the agent use YOUR local model instead of the cloud.

Key concepts:

Model selector, bottom-right corner.

What it is:

Hermes only works well with the 64k model; that’s why you prepared the derived version beforehand.

Why learn:

Closes the loop: module 2.4 exists exactly for this moment.

Key concepts:

Context requirements, the right model.

What it is:

hermes doctor checks the installation's health, and hermes status shows its current state.

Why learn:

They’re your first commands when something won’t connect.

Key concepts:

hermes doctor, hermes status.

What it is:

Send a "hi" in Hermes with the local model selected and see the response come from your machine.

Why learn:

It’s the final proof that the agent is running 100% locally.

Key concepts:

Smoke test, local response.

View Full
2.6~30 min

🖥️ Desktop app, terminal, and Telegram

Three ways to communicate with Hermes: the desktop app (user-friendly), the terminal (powerful), and Telegram (from anywhere) — plus sessions, branch/fork, and artifacts.

0 of 6 · 0%
What it is:

A Hermes app window, friendlier than the terminal for beginners.

Why learn:

It’s the path of least friction for using the agent day to day.

Key concepts:

Desktop app, graphical interface.

What it is:

From the terminal, you can run hermes dashboard, hermes setup, and the agent’s other commands.

Why learn:

The terminal is the most powerful and automatable route.

Key concepts:

hermes dashboard, hermes setup.

What it is:

You can chat with the agent on Telegram, like messaging a contact.

Why learn:

And that’s what lets you use your phone’s agent from anywhere (Project 7 of Track 3).

Key concepts:

Telegram, remote access, chat.

What it is:

Each conversation is a “session”; you open New session, and Hermes saves the history.

Why learn:

Organizing by sessions keeps contexts separate (work, study, etc.).

Key concepts:

Session, New session, history.

What it is:

You can "fork" a conversation at a point and continue along two different paths.

Why learn:

Lets you test approaches without losing the original thread.

Key concepts:

Branch, fork, parallel lines.

What it is:

The agent's outputs (code, text, files) can be saved and revisited as artifacts.

Why learn:

And it’s where the agent’s work “lives,” ready to use later.

Key concepts:

Artifacts, saved outputs, reuse.

View Full