PTENES
MODULE 3.1 · PROJECT

💬 Project 1: 100% local private chat

Enough theory. This is your first end-to-end project: build an AI chat that runs entirely on your machine, responds with the internet turned off, and even lets you fork the conversation into two lines of reasoning. Follow the steps — each one has an objective, instructions, how to verify it, and the expected result.

6
Steps
~35
Minutes
Practical
Level
Project
Type
⬇️pull model 💬open chat ✈️test offline 🌿fork ✅private chat

The project path in one line: from model pull until the private chat, through the test offline — the moment that proves everything runs on your machine.

1

🎯 Goal: chat 100% offline

The goal of this project is simple and powerful: to have a AI chat that works without internet, running entirely on your computer. No messages leave the machine, there's no cost meter, and AI remains available even offline. It's the clearest proof of everything you saw in Tracks 1 and 2.

🧭 What you’ll have at the end

  • •An open model downloaded and ready on your machine.
  • •A chat running in the terminal or in the Ollama app.
  • •Living proof that it responds with Wi-Fi turned off.
  • •A conversation forked into two lines of reasoning.

New here? "Offline" means without any connection to the internet. Since the model and your data are on your disk, the chat doesn’t need to ask a server for anything—that’s why it works on a plane, in the field, or when the network goes down.

Key concepts

Local chat

The conversation happens on your machine, from start to finish.

Offline

Works without internet — the AI is on your disk.

$0 per use

After downloading, every conversation is free.

Privacy

Nothing you write leaves the machine.

2

⬇️ Make sure a model is installed on your machine

Step objective: have the course’s fast model downloaded to disk. Without a local model, there’s nothing to run offline. You’ll use the Qwen 3 30B-A3B in the quantized (lighter) version, which is the "fast model" shown in the video.

🎯 Step by step

  1. Open the terminal (or use the Ollama app, installed in Track 2).
  2. Run the pull command below to download the model (≈18 GB).
  3. Wait for the progress bar to reach 100%.

Command to paste (goal: download the model):

ollama pull qwen3:30b-a3b-q4_K_M

How to verify (the model appears in the list):

ollama list
# saida esperada (exemplo):
# NAME                       SIZE
# qwen3:30b-a3b-q4_K_M       18 GB
Conversation with the Qwen model running locally: the model shows 'Thought for 6.2 seconds' before responding with facts about color theory
Video frame: after the pull, this is the model you’ll talk to. Notice "Thought for 6.2 seconds" — the model "thinks" before responding; that’s normal, and you’ll see it in Module 3.1, step 5.

💡 Practical tip

The passage q4_K_M and the quantization level (the model "shrunk" to fit better in memory). If your machine has little RAM, it's worth downloading a smaller model first, testing it, and deleting it with ollama rm <modelo>. Exploring is cheap.

Key concepts

ollama pull

Downloads the model to your disk.

Quantization

q4_K_M makes the model smaller and lighter.

ollama list

Confirm that the model is on the machine.

One-time download

Download once; then run offline.

3

💬 Open the chat (terminal or app)

Step objective: start the conversation. You have two ways to use the same model — the terminal (fast, direct) and the Ollama app (with a chat window). Choose whichever feels more comfortable; both use exactly the same local model.

✓ Terminal path

  • ✓Fast and always available.
  • ✓Works on any OS.
  • ✓Exit the chat with /bye.

🖼️ App route

  • •Friendly chat window.
  • •Good for people who don’t like the terminal.
  • •Choose the model from the app’s list.

Command to paste (goal: open the chat in the terminal):

ollama run qwen3:30b-a3b-q4_K_M
# o prompt fica aguardando voce digitar.
# experimente: "explique em 2 linhas o que e teoria das cores"
# para sair do chat, digite: /bye

How to verify: the model answers your first question. The first time you use it, it may take a little longer (while loading into memory)—that’s normal.

New here? ollama run does two things: if the model isn't loaded, it loads it into memory; then it opens a chat right in the terminal. You can use run directly without having done the pull before—it downloads automatically—but separating the steps makes the process clearer.

Key concepts

ollama run

Loads the model and opens the chat in the terminal.

/bye

Exit the terminal chat.

Ollama app

The same model, with a graphical interface.

1st load

The first response takes longer (the model is loading).

4

✈️ Test offline (the moment of truth)

Step objective: prove, beyond any doubt, that the chat doesn’t depend on the cloud. You will turn off the internet and ask another question. If the model responds, that's proof: the intelligence is on your machine.

1

With the chat open, turn off Wi-Fi

Disconnect from the network (turn off Wi-Fi or unplug the cable). In the video, it literally means “pull the cable.”

2

Ask a new question

Ask anything. For example: “summarize color theory in 3 points.”

3

Watch the response arrive

The model responds normally — without a network connection. Test complete.

CLOUD · offline you serverunreachable LOCAL · no internet you local modelresponds ✓

On the left, with no network, the call to the server fails; on the right, the local model works even offline — exactly what you just tested.

Why this matters: this test is what turns "I believe it’s local" into "I saw it work". In flight, somewhere without a signal, or when the network keeps dropping, your chat keeps working.

Key concepts

Offline test

Proof that nothing depends on the cloud.

Availability

The AI is always there, without relying on a network.

Air-gapped

Without a network connection, there’s no way for the data to get out.

Resilience

A network outage doesn’t leave you stuck.

5

🌿 Fork the conversation into two lines

Step objective: learn to create a fork (branch) off a conversation to explore two paths from the same point without losing the original. In the Hermes app, this appears in the Sessions list, with conversations that branch off.

Hermes screen with the Sessions list and pinned conversations (Pinned), showing a conversation branching into two lines (branch/fork)
Video frame: notice the Sessions list on the left. Forking means taking a message and opening a second path from it — you keep the original conversation intact and create an "alternate version" to compare responses.

🎯 How to do it (action in the interface)

  1. Have an ongoing conversation.
  2. Choose the message you want to branch from.
  3. Create its fork/branch (the UI shows the resulting Sessions).
  4. Ask a different question in each branch and compare.

How to verify: two conversation branches appear side by side in Sessions, starting from the same message.

📊 When forking helps

  • •Test two response tones (formal vs. informal) without redoing everything.
  • •Explore two solutions to the same problem in parallel.
  • •Keep the original conversation as the "source of truth".

New here? "Fork" (or "branch") comes from the world of code: creating a branch from a common point. In a conversation, it means opening a copy that follows a different path—without deleting the original.

Key concepts

Fork / branch

A branch of the conversation starting from a shared point.

Sessions

The list of conversations, including branched ones.

Compare paths

Two responses from the same message.

Original conversation

Remains intact—it isn’t deleted by the fork.

6

✅ Result: private chat without internet

Expected result: you have an AI chat that is your — private, free, and works anywhere. It’s the course’s first “this is really mine” moment, and the foundation that Project 2 will build on to create the complete Hermes agent in Vault.

🧾 Completion checklist

✓ollama list shows the downloaded model.
✓The chat responded through the terminal (ollama run) or through the app.
✓It answered again with the internet turned off.
✓You forked the conversation into two paths.

⚠️ Did it go wrong?

  • ✗"Command not found": reopen the terminal or check the PATH (covered in Module 2.1).
  • ✗Very slow or freezing: the model may be too large for your RAM—try a smaller one.
  • ✗"Doesn’t respond offline": make sure the pull is finished (100%) before disconnecting the network.

Optional self-check: What’s the PROOF that the chat is 100% local?

🎯 Project summary

✓
Objective achieved — a 100% local, private, and free AI chat.
✓
Model + chat — ollama pull e ollama run bring the model and open the conversation.
✓
Offline proof — it answered with the network off: the intelligence is on your machine.
✓
Conversation fork — you explore two paths without losing the original.

Next project:

3.2 — Hermes agent running locally (Vault end to end): from chat to a full agent, 100% private.