🏠 The model on your machine
Note the difference from module 1.4: there, the Hermes ran locally. Here, the AI MODEL itself runs on your machine — nothing needs to leave your computer.
🧬 What changes
- •Before: the brain was on a remote server (API/OAuth)
- •Now: the brain runs on your machine
- •Result: complete privacy and offline operation
💡 Practical tip
"Local-hosted" can mean two levels: Hermes running locally (1.4) and the local model (this module). The latter is what ensures absolute privacy.
⚖️ The trade-off: limited by hardware
Large models have billions of parameters and need a data center. Locally, you are limited by your hardware — privacy costs performance and speed.
✓ What you gain
- ✓100% privacy — nothing leaves the machine
- ✓Works offline
- ✓No per-token cost
✗ What you give up
- ✗Less performance than huge servers
- ✗Less speed
- ✗Cap set by your hardware
📐 How to Size It
Don’t guess what fits. Go to Apple → About This Mac, send Hermes a screenshot and ask: "what's the most powerful local model I can run?".
Open “About This Mac”
Apple menu → About This Mac. It shows memory, chip, and specs.
Send the screenshot to Hermes
It reads the specs visually and understands what your machine can handle.
Ask which model is ideal
"What’s the most powerful local model I can run?" — it recommends one.
💡 Practical tip
Letting Hermes read your specs prevents the mistake of downloading a model that’s too large and freezes the machine.
🦙 Download via Ollama
The default tool for running models locally is OllamaExamples: Gemma, Qwen 32B/3.6, or free Cloud options.
Download and run (illustrative)
$ ollama pull gemma # modelo pequeno, leve $ ollama pull qwen:32b # mais potente, exige mais RAM $ ollama run gemma # conversa local, offline
📊 Common options
- •Gemma — lightweight, a good starting point
- •Qwen 32B / 3.6 — more powerful, requires more hardware
- •Cloud free — when the local setup can’t handle it
🔒 100% private and offline
Since everything runs locally, it’s 100% private and it works offline: 1000 ft underground, flying in an airplane, or even in space.
🛡️ The big argument for local
Nothing leaves your machine. For anyone handling sensitive data or needing to work offline, that’s the difference between being able to use AI and not being able to.
- •No internet dependency
- •No data sent to external servers
🧭 When it's worth it
Local makes sense when privacy or offline operation are nonnegotiableFor maximum power, cloud models still come out ahead — it’s a choice based on priorities.
✓ Prefer local when
- ✓Privacy is nonnegotiable
- ✓Need to work without internet
- ✓Want zero cost per use
✗ Prefer the cloud when
- ✗Need maximum power/speed
- ✗Your hardware is limited
- ✗The task requires the best possible model
Wrap up Track 1: you already know what an agent is, what Hermes is, where it lives, how to connect models, and how to run everything locally. Track 2 covers capabilities (memory, soul, MCPs).
🧯 Common mistakes when running locally
People who test local models often run into the same problems. Anticipate them so you don't bog down your machine or get frustrated with performance.
✓ Do it
- ✓Check the specs first (send the screenshot)
- ✓Start with a lightweight model (Gemma) and move up
- ✓Use local when privacy/offline access are priorities
✗ Avoid
- ✗Download a model that's too large and freeze everything
- ✗Expecting datacenter performance on a laptop
- ✗Insisting on local processing when the task requires power
Pocket decision (illustrative)
privacidade inegociável / offline -> 🔒 modelo LOCAL (Ollama) máxima potência / hardware fraco -> 🌐 modelo em NUVEM (API/OAuth) não sei o que cabe -> mande specs ao Hermes e pergunte
📌 Module Summary
Next Track:
Track 2 — Capabilities (memory, soul, integrations, MCPs)