🛠️ Hands-on
Enough theory — time to install. You’ll install Ollama on your computer, choose a model that fits your hardware, download it and chat with it, prepare the agent model (Qwen 3 Coder with 64k context), and connect everything to the Hermes Agent — desktop, terminal, or Telegram. Each module includes commands ready to copy and paste.
The four practical steps in this track, from left to right: install Ollama, download a model, create the 64k version, and plug everything into the Hermes Agent — arriving at an agent running 100% on your machine.
Track map
⬇️ Install Ollama
One command and you’re done
🎚️ The right model
Fit your hardware
💬 Your 1st model
Download and chat
🪟 The agent model
Qwen 3 Coder 64k
🔌 Connect to Hermes
Plugged-in local model
🖥️ Desktop and Telegram
Three ways to communicate
Detailed content
⬇️ Install Ollama
The first practical step: get Ollama from the website, install it via the terminal (Mac/Linux) or the installer (Windows), and check that everything worked — with commands ready to paste.
You start on the official ollama.com site, where you’ll find the Download button and the installation command.
Getting it from the official source avoids fake versions and ensures you get the right installer for your system.
Official source, Download, system detection.
On Mac/Linux, a single command (curl ... | sh) downloads and installs Ollama all at once.
It’s the fastest way and the one you’ll actually paste in—without clicking through screens.
curl, pipe to sh, installation script.
On Windows (and Mac too), you can download the installer from the website and click to install.
If you don’t like the terminal, this is the command-line-free path.
Graphical installer, tray app.
Running ollama --version and ollama list confirms it’s installed and responding.
Checking before moving on prevents wasting time later on a "command not found" error.
Verification, --version, empty list.
Ollama has a windowed app and terminal commands—they do the same thing.
Knowing both exist lets you choose whichever feels more comfortable.
GUI vs CLI, same engine underneath.
The most common issue is that the terminal can't "see" the newly installed ollama—just reopen the terminal.
Resolves 90% of installation scares without needing to reinstall anything.
PATH, reopen the terminal, new session.
🎚️ Choose the right model for your hardware
Before downloading gigabytes, find out what your machine can handle: check the hardware, ask Hermes for a recommendation, understand the effort levels, and learn the rule of leaving RAM headroom.
Check how much memory (RAM) and which chip you have — on a Mac, in “About This Mac.”
RAM determines which model fits; without that, you're downloading blindly.
RAM, chip/GPU, “About This Mac”.
You show Hermes your hardware, and it suggests models that will fit (e.g., M4 Max 36GB → Qwen 32B / 30B-A3B).
Takes the guesswork out of the choice; the agent itself points you in the right direction.
Recommendations by hardware, options by size.
In the Hermes selector, you adjust the effort (Minimal, Low, Medium, High, Max) and toggle Thinking/Fast.
This is how you balance quality and speed without switching models.
Effort, Thinking, Fast.
The model needs to fit in RAM with room for the rest of the system; if it’s too full, the machine freezes.
Avoids the classic mistake of choosing the largest model and making the machine choke.
Headroom, RAM buffer, download→test→delete.
The q4_K_M suffix indicates a quantized (compressed) model that takes up less memory.
Understanding quantization lets you fit larger models in the same RAM.
Quantization, q4_K_M, size vs. accuracy.
A concrete recommendation to get started: qwen3:30b-a3b-q4_K_M, fast and well-balanced.
Starting with a tested model helps you avoid choice paralysis.
First model, qwen3:30b-a3b-q4_K_M.
💬 Download and chat with your 1st model
The “wow” moment: download the model, run it in the terminal or app, understand “thinking,” and manage what’s on disk — all running offline on your machine.
The ollama pull command downloads the chosen model (about 18 GB) to your drive, with a progress bar.
And it’s the only time you need the internet; after that, it runs offline.
pull, one-time download, progress bar.
ollama run opens a chat in the terminal itself; you type, it responds, and you exit with /bye.
It’s the most direct way to prove that the local model works.
run, prompt in the terminal, /bye.
The Ollama app has a chat window; choose a model from the list and chat as you would in a regular app.
It’s the comfortable path for those who prefer the mouse to the keyboard.
Chat app, model selector.
Reasoning models "think" before responding; the app shows something like "Thought for 6.2 seconds".
Explains why the response takes a little while — and why it’s usually better.
Thinking, reasoning, “thinking” time.
The first time you ask a question, the model loads into memory and takes longer; after that, it gets faster.
Knowing this helps you avoid thinking it’s “frozen” right away.
Cold start, loading into RAM, warm-up.
ollama list shows what you downloaded, ollama ps shows what’s running, and ollama rm deletes a model.
You’ll try several; deleting the ones you don’t use frees up disk space.
list, ps, rm, disk cleanup.
🪟 The agent model: Qwen 3 Coder 64k
The agent requires 64k of context. You’ll understand why, learn about Qwen 3 Coder, write a Modelfile that raises num_ctx, and create the 64k derived model with one command.
Hermes Agent needs a model with 64,000 context tokens; the one from module 2.3 doesn’t always have that.
It’s why you prepare a specific model before connecting the agent.
64k requirement, agent model.
An open model from the Qwen family designed for coding and tool use—ideal for an agent.
Knowing why it’s chosen helps you switch later with confidence.
Qwen 3 Coder, a model for agents and tools.
A Modelfile is a short recipe: it starts with a base model and sets num_ctx to 65536 (64k).
This is how you “make” the 64k model from one you already have.
Modelfile, FROM, PARAMETER num_ctx 65536.
ollama create qwen3-coder-64k -f Modelfile creates a new model with the context set to 64k.
This derived model is what you’ll point Hermes to in module 2.5.
create, derived model, -f Modelfile.
ollama show qwen3-coder-64k and ollama list confirm that the model exists and has 64k.
Checking before plugging it into the agent prevents context errors down the line.
show, list, check num_ctx.
A larger context uses more RAM; 64k uses more than the same model's default context.
Turn the headroom rule back on: you need spare capacity to run the agent.
Context memory cost, headroom.
🔌 Connect the local model to the Hermes Agent
The final connection: install/update Hermes (open source, MIT, Nous Research), select the local 64k model, diagnose and test the connection while running 100% offline.
Hermes Agent is the open-source “AI OS” from Nous Research, under the MIT license.
Open source + MIT = you can run, audit, and adapt it without restrictions.
Open source, MIT license, Nous Research.
hermes update brings Hermes up to the latest version; when it finishes, "HERMES IS READY" appears.
Starting with the latest version helps you avoid bugs that have already been fixed.
hermes update, "HERMES IS READY".
In the Hermes selector, you choose qwen3-coder-64k; the active model appears in the lower-right corner.
It’s the step that makes the agent use YOUR local model instead of the cloud.
Model selector, bottom-right corner.
Hermes only works well with the 64k model; that’s why you prepared the derived version beforehand.
Closes the loop: module 2.4 exists exactly for this moment.
Context requirements, the right model.
hermes doctor checks the installation's health, and hermes status shows its current state.
They’re your first commands when something won’t connect.
hermes doctor, hermes status.
Send a "hi" in Hermes with the local model selected and see the response come from your machine.
It’s the final proof that the agent is running 100% locally.
Smoke test, local response.
🖥️ Desktop app, terminal, and Telegram
Three ways to communicate with Hermes: the desktop app (user-friendly), the terminal (powerful), and Telegram (from anywhere) — plus sessions, branch/fork, and artifacts.
A Hermes app window, friendlier than the terminal for beginners.
It’s the path of least friction for using the agent day to day.
Desktop app, graphical interface.
From the terminal, you can run hermes dashboard, hermes setup, and the agent’s other commands.
The terminal is the most powerful and automatable route.
hermes dashboard, hermes setup.
You can chat with the agent on Telegram, like messaging a contact.
And that’s what lets you use your phone’s agent from anywhere (Project 7 of Track 3).
Telegram, remote access, chat.
Each conversation is a “session”; you open New session, and Hermes saves the history.
Organizing by sessions keeps contexts separate (work, study, etc.).
Session, New session, history.
You can "fork" a conversation at a point and continue along two different paths.
Lets you test approaches without losing the original thread.
Branch, fork, parallel lines.
The agent's outputs (code, text, files) can be saved and revisited as artifacts.
And it’s where the agent’s work “lives,” ready to use later.
Artifacts, saved outputs, reuse.