PTENES
MODULE 1.7

💻 Local & Private

Run the MODEL on your machine, not just Hermes. The trade-off: you’re limited by the hardware, but you get complete privacy and offline operation — 1000 ft underground, in flight, or in space.

Gemma (small)hardware: light Qwen 32Bhardware: medium the largest that fitsYOUR hardware limit giant modelsneed a data center 🔒 private + offline
6
Topics
24
Minutes
Basic
Level
Practical
Type
1

🏠 The model on your machine

Note the difference from module 1.4: there, the Hermes ran locally. Here, the AI MODEL itself runs on your machine — nothing needs to leave your computer.

🧬 What changes

  • •Before: the brain was on a remote server (API/OAuth)
  • •Now: the brain runs on your machine
  • •Result: complete privacy and offline operation

💡 Practical tip

"Local-hosted" can mean two levels: Hermes running locally (1.4) and the local model (this module). The latter is what ensures absolute privacy.

2

⚖️ The trade-off: limited by hardware

Large models have billions of parameters and need a data center. Locally, you are limited by your hardware — privacy costs performance and speed.

✓ What you gain

  • ✓100% privacy — nothing leaves the machine
  • ✓Works offline
  • ✓No per-token cost

✗ What you give up

  • ✗Less performance than huge servers
  • ✗Less speed
  • ✗Cap set by your hardware
3

📐 How to Size It

Don’t guess what fits. Go to Apple → About This Mac, send Hermes a screenshot and ask: "what's the most powerful local model I can run?".

1

Open “About This Mac”

Apple menu → About This Mac. It shows memory, chip, and specs.

2

Send the screenshot to Hermes

It reads the specs visually and understands what your machine can handle.

3

Ask which model is ideal

"What’s the most powerful local model I can run?" — it recommends one.

💡 Practical tip

Letting Hermes read your specs prevents the mistake of downloading a model that’s too large and freezes the machine.

4

🦙 Download via Ollama

The default tool for running models locally is OllamaExamples: Gemma, Qwen 32B/3.6, or free Cloud options.

Download and run (illustrative)

$ ollama pull gemma       # modelo pequeno, leve
$ ollama pull qwen:32b    # mais potente, exige mais RAM
$ ollama run gemma        # conversa local, offline

📊 Common options

  • •Gemma — lightweight, a good starting point
  • •Qwen 32B / 3.6 — more powerful, requires more hardware
  • •Cloud free — when the local setup can’t handle it
5

🔒 100% private and offline

Since everything runs locally, it’s 100% private and it works offline: 1000 ft underground, flying in an airplane, or even in space.

⛏️
1000 ft below
✈️
Flying
🚀
In the space

🛡️ The big argument for local

Nothing leaves your machine. For anyone handling sensitive data or needing to work offline, that’s the difference between being able to use AI and not being able to.

  • •No internet dependency
  • •No data sent to external servers
6

🧭 When it's worth it

Local makes sense when privacy or offline operation are nonnegotiableFor maximum power, cloud models still come out ahead — it’s a choice based on priorities.

✓ Prefer local when

  • ✓Privacy is nonnegotiable
  • ✓Need to work without internet
  • ✓Want zero cost per use

✗ Prefer the cloud when

  • ✗Need maximum power/speed
  • ✗Your hardware is limited
  • ✗The task requires the best possible model

Wrap up Track 1: you already know what an agent is, what Hermes is, where it lives, how to connect models, and how to run everything locally. Track 2 covers capabilities (memory, soul, MCPs).

7

🧯 Common mistakes when running locally

People who test local models often run into the same problems. Anticipate them so you don't bog down your machine or get frustrated with performance.

✓ Do it

  • ✓Check the specs first (send the screenshot)
  • ✓Start with a lightweight model (Gemma) and move up
  • ✓Use local when privacy/offline access are priorities

✗ Avoid

  • ✗Download a model that's too large and freeze everything
  • ✗Expecting datacenter performance on a laptop
  • ✗Insisting on local processing when the task requires power

Pocket decision (illustrative)

privacidade inegociável / offline  -> 🔒 modelo LOCAL (Ollama)
máxima potência / hardware fraco   -> 🌐 modelo em NUVEM (API/OAuth)
não sei o que cabe                 -> mande specs ao Hermes e pergunte

📌 Module Summary

✓
Local model — not just Hermes; the model itself runs on the machine.
✓
Trade-off — limited by the hardware; lower performance/speed.
✓
Scale — screenshot of "About This Mac" → ask Hermes.
✓
Ollama — Gemma, Qwen 32B/3.6, free cloud.
✓
Private & offline — 1000 ft underground, flying, or in space.

Next Track:

Track 2 — Capabilities (memory, soul, integrations, MCPs)