Trail map
🤖 Agent, not chatbot
Does, doesn't just talk
🐕 What Hermes is
Your dog, not the contractor
🧠 One Brain, 22 Mouths
One mind, everywhere
🏠 Where Hermes Lives
Local, VPS, or cloud
🔑 OAuth vs. API key
Button that comes back vs. key
🎛️ Choosing the Model
Toolbox, not a hammer
💻 Local & Private
1000 ft below ground
Detailed content
🤖 Agent, not chatbot
The difference that changes everything: a chatbot tells you how to do something; an agent does it. You give it a goal, it has tools, and it executes.
A chatbot explains HOW to book a flight; an agent actually books it. The difference is the ability to act in the world.
It's the core concept of all of Track 1. Without it, you use Hermes like just another chatbot and waste its power.
Talks vs. does; advice vs. execution; response vs. completed action.
Chatbot = a smart friend who gives tips. Agent = a personal assistant who does the task for you.
The analogy makes the practical difference clearer than any technical definition.
Advice vs. delegation; you stay in command, it handles execution.
"Find the cheapest flight from Dubai to Toronto in the next 2 weeks" → it searches, picks one, and opens the result in a nice-looking HTML page.
Shows the goal → research → decision → delivery cycle that defines an agent.
One goal; multiple autonomous steps; ready-to-use result.
The agent has access to tools (Gmail, calendar, search, browser), and it acts through them rather than just chatting.
Tools are what turn text into real action—the essence of an agent.
No tools = chatbot; with tools = agent.
Hermes is an AI with tools, accessible wherever you are through your phone.
Mobile availability is what makes the agent useful day to day, not just at your desk.
Mobile access; an agent always at hand.
For a quick question ("what does X mean?"), a chatbot is enough; an agent shines when there's a TASK to carry out.
Avoids using (and paying for) a heavy tool when a simple answer will do.
Question = chatbot; task with steps = agent.
🐕 What Hermes is (and when to use it)
Where Hermes fits vs. Claude Code, OpenClaw, and IDEs. It’s your Labrador: it lives with you, knows you, and gets better year after year.
Hermes is like your dog: it lives with you, knows you, and gets better year after year.
Defines Hermes’s “design ethos”: persistence and growth over time.
Persistence; long-term relationship; continuous improvement.
Claude Code is the contractor: excellent for a specific job, but it doesn't live with you or remind you later.
Knowing the difference prevents you from expecting a session tool to have persistent memory.
Session vs. persistence; precision tool vs. companion.
OpenClaw is your roommate; anti-gravity is your IDE/coding partner. Each has a different role.
Mapping the ecosystem helps you choose the right tool for each moment.
Distinct roles; neither replaces the other.
The essence of Hermes: the more you use it, the more it learns about you and the more useful it becomes.
Sets the right expectation: the value grows over time, not on day one.
Accumulation effect; value compounds with use.
Use Hermes when you're on the go — at a coffee shop, at the gym, on your phone. Use desktop tools when you're at your desk.
Knowing the right context for each tool maximizes productivity.
Mobile = Hermes; desktop = desktop tools.
Hermes’s design principle is to live with you and improve over time—not be a disposable tool.
This ethos explains decisions like local memory, soul, and backups.
Persistent companion; growth; relationship.
🧠 One Brain, 22 Mouths
The same intelligence, accessible through 22+ interfaces. Where Hermes runs is the brain; the channels are just mouths plugged into it.
Where Hermes runs (your PC or a VPS) is the brain. It's the SAME intelligence, no matter how you talk to it.
Understanding that the brain is unique explains why memory and context are shared across channels.
A central brain; pluggable interfaces.
Telegram, Discord, WhatsApp, Slack, Matrix, browser, your own OS — 22+ interfaces for the same mind.
You talk to the agent through the channel you already use, without learning a new app.
Multichannel; no required app.
Imagine a middle layer: it receives calls from any channel and routes everything to your agent.
The analogy explains how messages from different sources reach the same brain.
Central routing; channel convergence.
Since intelligence is central, you’re not tied to a single software or interface.
Reduces vendor lock-in and gives you channel flexibility based on context.
No lock-in; switch channels freely.
Since there’s one brain, you can start a conversation on Telegram and continue in the browser without losing the thread.
It's the most immediate practical benefit of having a central brain.
Continuity across channels; unified memory.
Besides chat apps, Hermes has its own OS/dashboard as an interface — another mouth for the same brain.
Introduces the concept of an Operating System, explored in more depth in Track 3.
OS as a channel; unified view.
🏠 Where Hermes Lives
Terminal, local, or VPS: the three possible homes for Hermes, with the cost, control, and security trade-offs of each.
A terminal is where you enter commands; local means your own computer; a VPS is a computer owned by another company that you rent.
Without this vocabulary, the rest of the module (and the installation) won't make sense.
Terminal; local; VPS.
Running it on your PC is free, easy, and secure; it works as long as the computer is on. Many people use an old MacBook 24/7.
It’s the recommended way to get started—zero hosting costs and complete control.
Free; must be powered on; dedicated machine 24/7.
You install it by pasting the command into the terminal (Cmd+Space → "terminal" → paste) or asking Claude Code to install it for you.
It's the practical step that takes Hermes off the page and gets it running.
Command pasted; Claude Code as the installer.
A VPS runs on another company's computer (hosting): you rent it, pay monthly, and need to protect the ports to prevent attacks.
It’s the option for keeping Hermes running 24/7 without leaving your personal PC on.
Monthly rent; uptime; attack surface.
VPS providers often pay referral fees through affiliate links. Not every recommendation is neutral.
Makes you more discerning when choosing a provider and helps avoid biased decisions.
Affiliate links; commercial bias.
Since it runs on your machine, you can ask Hermes to "shut down the gateway" or "restart"—it controls its own infrastructure.
Shows the power of local control: the agent manages itself.
Self-administration; control through natural language.
🔑 OAuth vs. API key
The two ways to connect a model to Hermes. OAuth is the button you can revoke; the API key is the key you keep and can rotate.
With OAuth, the browser opens, you log in and click “Allow.” The connection is ready, with no need to handle keys.
It’s the simplest and safest way to connect when the provider offers it.
Login + Allow; no exposed key; revocable.
An API key is a string of characters stored on a server that grants access to the provider’s models.
It’s the method used when OAuth isn’t available (e.g., Claude) and provides broad access.
Secret string; access to all of the provider's models.
You can rotate the API key at any time; once rotated, the old key will never work again.
It's your defense if a key leaks — just rotate it.
Rotation; immediately revoke the old key.
Grok and ChatGPT connect via OAuth; Claude does NOT offer OAuth—it’s API key only.
Knowing who offers what prevents you from looking for an OAuth button that doesn't exist.
OAuth: Grok, ChatGPT; API key: Claude and others.
You run homes setup in the terminal, choose the provider and then complete OAuth (reauthenticate) or paste the API key.
It's the exact point where the two methods meet in practice.
homes setup; choose provider; OAuth or paste key.
OAuth is the button you can "take back" at any time; an API key is a key that needs to be stored securely and can be rotated.
The image illustrates the difference in responsibility between the two methods.
Revoke vs. keep; convenience vs. control.
🎛️ Choosing the Model
Hermes is an agentic framework; the brain is a swappable model. Multi-brain strategy: the best model for each task.
Hermes is the agentic framework; the "brain" is a pluggable model. The command /model switches this brain.
Separating the framework from the model is what enables a model-agnostic strategy.
Fixed framework; swappable model; /model.
For heavy reasoning, use Opus 4.7/4.8 via OpenRouter, with a spending cap (e.g., US$10/month, then stop).
Expensive model only where it’s worthwhile, with a limit to keep costs from getting out of hand.
Opus for reasoning; spending cap.
For high-volume general tasks, use GPT (your ChatGPT via OAuth makes use of the US$20 subscription) or Grok (with X connected, Twitter search).
Makes use of subscriptions you already pay for and reduces cost per token.
GPT via OAuth; Grok + X; low-cost volume.
DeepSeek and free models run at almost no cost. For example: DeepSeek V4 flash delivers ~95% of the performance for ~1% of the cost.
For autopilot and background tasks, a cheap model is the obvious choice.
~95% of the performance, ~1% of the cost; free for high-volume use.
OpenRouter is one connection that gives you access to hundreds of models, with performance and cost rankings to compare.
Centralizes access and makes it easy to switch models based on the task.
Single hub; rankings; cost comparison.
Using a single model for everything is like having only one hammer: everything becomes a nail. The idea is to vary the tool by task.
It's the practical summary of the multi-brain strategy.
Model-agnostic; the right tool for each job.
💻 Local & Private
Run the MODEL on your machine, not just Hermes. 100% private, works offline — 1000 ft underground, in flight, or in space.
Local-hosted here means that not only Hermes, but the AI MODEL itself, runs on your machine.
It's what guarantees total privacy and offline operation.
Local model; no external server.
Large models have billions of parameters and need a data center. Locally, you’re limited by your hardware.
Sets expectations: privacy costs performance/speed.
Parameters vs. hardware; lower performance.
Go to Apple → About This Mac, send Hermes a screenshot, and ask “what’s the most powerful local model I can run?”
It's the practical way to find out what your hardware can handle without guessing.
Hardware specs; recommendation based on a screenshot.
You download models via Ollama — e.g., Gemma, Qwen 32B/3.6, or free Cloud options.
It’s the go-to tool for running models locally with just a few commands.
Ollama; Gemma; Qwen.
Since everything runs locally, it’s 100% private and works offline: 1000 ft underground, flying, or in space.
It's the main argument for local: nothing leaves your machine.
Complete privacy; offline; no network dependency.
Local is a good choice when privacy or offline operation is nonnegotiable; for maximum power, cloud models still win.
Helps you choose between privacy and performance depending on the situation.
Priority-based decision; privacy vs. power.