PTENES
MODULE 3.2 · PROJECT

🤖 Project 2: Hermes agent running locally (Vault end-to-end)

The leap from Project 1: from chat to full agent. You’ll connect the 64k model to the Hermes Agent, turn on Vault mode (airgapped), and have the agent perform a real task — all without a single byte leaving the machine. This is the course’s "holy grail": capable, 100% private, and $0.

6
Steps
~40
Minutes
Practical
Level
Project
Type
🪟64k model 🔗connect 🗄️Vault on 🧪real task ✅nothing came out

The project’s path: from 64k model until confirmation that nothing came out, with the Vault connected in the middle — the step that makes the agent safe for sensitive data.

1

🎯 Goal: the entire agent in the Vault

The goal of this project is to run the Hermes Agent from start to finish using only your local model, in Vault mode — isolated from the internet. Unlike Project 1 (which was a chat), here the agent acts: it uses tools and carries out tasks, but everything stays on your machine.

Nous Research’s Hermes Agent website: open source, MIT license, with the motto “THE AGENT GROWS YOU”
Video frame: Hermes Agent is open source, licensed under MIT, from Nous Research. "Open source" + "MIT" mean you can freely view and use the code — a perfect fit with the course’s ideas of ownership and privacy.

🧭 What you’ll have at the end

  • •The 64k model (from Module 2.4) connected to Hermes.
  • •The agent operating in Vault mode, offline.
  • •A real task completed 100% locally.
  • •Confirmation that no data left the machine.

New here? "End-to-end" means the entire workflow—from your request to the result—happens without leaving the machine. "Air-gapped" is the security term for a system physically isolated from the network; Vault mode simulates this by disconnecting the agent from the internet.

Key concepts

Agent

LLM + tools: it takes action, not just responds.

Hermes (MIT)

Open source from Nous Research; you own the workflow.

Vault

Airgapped mode: the agent isolated from the network.

End-to-end

From the request to the result, everything on your machine.

2

🪟 Getting qwen3-coder-64k ready

Step objective: confirm that the 64k context model you created in Module 2.4 is available in Ollama. The Hermes agent requires 64k of context — a regular chat model won’t work, because memory and tools take up a lot of space.

🎯 Check the model (and recreate it if needed)

Command to paste (goal: check whether the 64k model exists):

ollama list
# procure por:
# qwen3-coder-64k

If it doesn’t appear, recreate it from the Modelfile in Module 2.4 (goal: derive the model with 64k context):

# Modelfile (conteudo):
# FROM <qwen3-coder:30b>
# PARAMETER num_ctx 65536

ollama create qwen3-coder-64k -f Modelfile

How to verify: ollama list shows qwen3-coder-64k. The section <qwen3-coder:30b> and the variable part—the base model you chose.

📊 Why 64k, in numbers

  • •65,536 tokens of context (≈ 25-30 thousand words of "working memory").
  • •The agent uses context for memory, instructions, and tool results.
  • •Without enough context headroom, the agent "forgets" the beginning of the task along the way.

Note: a large context uses memory. A model with 64k uses more RAM than the same model with a small context. If your machine is under strain, this is where it can freeze—check the RAM headroom (covered in Module 1.4).

Key concepts

num_ctx 65536

The parameter that defines the 64k context.

ollama create

Derive the model from a Modelfile.

Agent requirements

Hermes requires 64k; regular chat won’t work.

Memory cost

A large context uses more RAM.

3

🔗 Connect the model to Hermes

Step objective: make sure Hermes is up to date and select your local model as the agent's brain. First, the actual update command; then, the action in the interface to point the agent to your model.

🎯 Update and check Hermes

Command to paste (goal: get Hermes ready):

hermes update
# ao terminar, o painel mostra: "HERMES IS READY / Launch Hermes"

How to verify (the agent is healthy):

hermes status
Hermes Agent desktop app on the 'New session' screen, ready to start a conversation with the agent
Video frame: the Hermes desktop app on "New session." This is where you select the model. The selected model appears in the bottom-right corner — and where you confirm that the agent is using your qwen3-coder-64k local, not the cloud.

🖱️ Interface action (select the local model)

  1. Open the Hermes model selector.
  2. Choose your Ollama model: qwen3-coder-64k.
  3. Confirm: the name appears in the bottom-right corner.

How to verify: the bottom-right corner shows your selected local model.

Honesty: selecting the model is an action in the interface, not a terminal command. The actual Hermes commands you use here are hermes update e hermes status; the rest is clicking the selector.

Key concepts

hermes update

Update Hermes ("HERMES IS READY").

hermes status

Shows the agent's health.

Model selector

Where you point the agent to the local model.

Bottom-right corner

Where the active model is confirmed.

4

🗄️ Turn on Vault (airgapped)

Step objective: activate mode Vault, which isolates the agent from the network — like unplugging the internet cable. This is the technical guarantee that nothing leaks: the agent simply has no way out.

Hermes modes diagram: Vault (air-gapped) switching to Connected, under the title “Toggle your privacy”
Video frame: the "Toggle your privacy" diagram. For this project, you stay at the Vault (airgapped, all local). Notice that privacy here is a switch: you choose when the agent can or can’t communicate with the outside world.
1

Select Vault mode

In the interface, choose Vault. The agent will use only the local model.

2

To be sure, disconnect from the network

Wi-Fi off or cable unplugged—the physical proof of the air gap.

3

Confirm the status

Hermes indicates that it’s in Vault / offline.

✓ With Vault, you GAIN

  • ✓Customer/health/IP data never leaves.
  • ✓Works offline (on a plane, off-grid).
  • ✓$0 per use, even if you run it a lot.

✗ With Vault, you GIVE UP

  • ✗No fresh web search.
  • ✗Without the cloud’s frontier model.
  • ✗The speed is determined by your machine.

No ideology: Vault isn’t for “everything.” It’s the right mode for sensitive data and offline use. When the task is difficult and there’s no sensitive data, you can switch to Connected or Cloud — that’s what Project 6 teaches.

Key concepts

Vault mode

The agent disconnected from the internet.

Air-gapped

Without a network connection, there’s no way for the data to get out.

Toggle your privacy

Privacy becomes a button you control.

Trade-off

Total privacy in exchange for cloud power.

5

🧪 Real task, 100% local

Step objective: give the agent a concrete task and watch it use tools to solve—proving that it doesn’t just chat, it acts. All of this in Vault, without a network connection.

🎯 Task suggestions (action in the agent chat)

  1. "Read this file and give me a summary in 5 points."
  2. "Organize these scattered notes into a task list."
  3. "Write an email draft based on these bullet points."

How to verify: the task finishes with a useful result, and you see the agent “working” (using tools), not just responding.

📊 Set expectations

The local model runs the SWE-bench benchmark at around 74 (the Qwen mentioned in the video, "runs on a laptop"), while the cloud frontier is around ~88. In other words: very capable, but not a frontier model. For everyday tasks in Vault, it gets the job done.

Speed depends on your machine; the first response may take longer (the model is loading into memory).

Tip: start with a task you know how to do yourself. That makes it easy to assess whether the agent got it right—and you build confidence before delegating bigger things.

Key concepts

Use tools

The agent takes action (reads, organizes, writes), not just talks.

Verifiable task

Start with something you know how to evaluate.

~74 on SWE-bench

Capable, but not the cloud boundary.

Local speed

Depends on the hardware; the first load takes longer.

6

✅ Confirm privacy: nothing left your device

Expected result: a capable agent that ran a real task without sending anything to the internet. With the network disconnected and Vault enabled, you have the confirmation that makes the agent usable with sensitive data — the project's goal.

🧾 Completion checklist

✓ollama list shows the qwen3-coder-64k.
✓hermes status indicates that the agent is healthy.
✓The local model appears in the lower-right corner.
✓The task ran without an internet connection, in Vault.

⚠️ Did it go wrong?

  • ✗Agent can’t find the model: confirm the selector and that Ollama is running.
  • ✗"Insufficient context": make sure you used the 64k model, not the regular chat model.
  • ✗Frozen/slow: a 64k context uses a lot of RAM — close other open tasks.

Optional self-check: Why does the agent need the qwen3-coder-64k, rather than an ordinary chat model?

🎯 Project summary

✓
Objective achieved — the entire Hermes Agent running locally, in Vault, for $0.
✓
Connected 64k model — o qwen3-coder-64k connected to Hermes via hermes update/status.
✓
Vault enabled — air-gapped: the agent has no way to leak data.
✓
Real task + privacy — the agent acted offline, and nothing left the machine.

Next project:

3.3 — Memory, personas, and skills: make the agent OS your own.