Trail map
🛡️ Security
Principle of least privilege
⭐ Northstar (Goals)
Doesn’t stop until it hits the goal
👥 Sub-agents
Parallel team
💓 Heartbeat / Cron
Works while you sleep
💰 Budget & Tokens
73% is fixed overhead
🌐 Operating System
The System Behind Everything
🔗 One Brain
Hermes + Claude Code
Detailed content
🛡️ Security · Principle of least access
The house rules: give the agent only what it needs, and never expose secrets in chat.
Unlike a chatbot, an agent can send email, modify files, and act on your computer. Real access requires real care.
Before enabling any new capability, you need to understand what could go wrong.
Don’t hand over the keys to the kingdom; enable capabilities gradually.
Give the agent only the permission the task requires — nothing more. This applies to Hermes and any system.
If something goes wrong, the damage is contained to a minimum.
Reading ≠ writing ≠ sending; start restrictive and expand as needed.
Enable "read email" now, but hold off on "send/reply to email" until you trust how it behaves.
A mistaken send is irreversible; reading is safe.
Leave irreversible actions for last; test in read-only mode first.
The chat is saved and indexed in memory; a secret pasted there gets stored, especially if there’s a daily backup.
Leaking the key gives access to your model and your money.
Use environment variables; if one leaks, rotate it immediately.
When the agent asks for access, you decide: allow it once, for the session, or block it entirely.
Each level trades convenience for control.
"Once" for the unknown; "never" for the dangerous.
The same principles protect you in any AI agent or system, not just Hermes.
The security habit is transferable and stays with you.
Minimum access + secrets outside the chat = secure foundation.
⭐ Northstar (Goals) · Doesn't stop until it hits the goal
Instead of question→answer, you set a direction and the agent works in turns until it reaches the goal.
You set a goal, and the entire conversation exists to achieve it, not to answer a one-off question.
It's how to turn a chat into a purposeful mini-project.
Goal > one-off answer; the agent pursues, it doesn't just react.
When you receive /goal, Hermes signs a budget of about 20 turns to reach a decision.
Knowing the limit helps you calibrate the size of the goal.
Asks questions, doesn't get stuck in a loop; converges within the budget.
For goals that are too big, the super goals skill generates 4 to 10 actions (yours and the agent's) with a progress bar.
Not every goal fits into 20 turns; super goals break it into steps.
Mission control on the dashboard; visible progress.
Commands to start, pause, and clear a goal at any time.
You stay in the driver’s seat even during an autonomous task.
Pausing saves state; clearing resets the goal.
Chief wickham loops are goals that last weeks, for planning your life, not just a short project.
Connects the short term (turns) to the long term (life).
Recurring review loop; a horizon of weeks.
Goals shine when there’s a decision to make or a short project to deliver—not for one-off questions.
Using goals for everything wastes turns.
Deciding A vs. B is a perfect use case; “what time is it” isn’t.
👥 Sub-agents · Parallel team
Hermes spins up several subagents with fresh context and delegates — 6 working together can do in 1h what would take 6h.
Instead of doing everything in a single thread, Hermes becomes an orchestrator that delegates work to helpers.
Large tasks become manageable when broken into several small ones.
Orchestrator + executors; divide and conquer.
Each sub-agent is created with its own clean context window, focused only on its subtask.
Lean context = better and cheaper answers.
Isolation prevents confusion between tasks.
4 to 6 agents research at the same time and deliver together at the end, instead of one after another.
Parallelism is the power-user move that compresses time.
Breadth > depth when tasks are independent.
"Research the best AI companies" spins up 2 sub-agents: one covers the US, the other the international market.
Shows how to divide a broad scope into clear slices.
Split by dimension (region, topic); merge at the end.
The Hermes co-founder runs 12 parallel instances every day to build Hermes itself (issues, dogfooding, kanban).
It’s proof that parallelism scales far beyond 2 or 3.
Dogfooding at scale; a fleet of agents.
Each agent can have a role: research, writing, design, scheduler — like roles on a real team.
Clear roles prevent two agents from doing the same thing.
Specialization; each role with its ideal model.
💓 Heartbeat / Cron · Works while you sleep
The heartbeat that keeps the agent alive 24/7 and the schedules that trigger tasks automatically.
A cron job that keeps the agent awake all the time, instead of just reacting when you talk to it.
It's what separates a chatbot from an agent that operates on its own.
Periodic heartbeat; continuous presence.
A sub-agent pings every few seconds to detect “zombie” (stalled) processes and recover them.
Without this, a stuck task would stay stuck forever.
Stall detection; fresh agent reclaim; retries.
You schedule an action to run at an interval: "remind me in 30s that the sky is blue".
Automates reminders and routines without you opening the app.
Time-based scheduling; runs on its own.
Every day at 8 a.m., Hermes scans email + calendar + what it knows about you and delivers 5 items that matter.
Shows how cron + memory become a proactive assistant.
Source aggregation; daily curation.
With a heartbeat, the agent proactively checks in: "periodically ask me things."
It reverses the relationship: it reaches out to you, instead of just responding.
Proactivity; agent initiative.
Each trigger uses tokens; too many scheduled tasks create costs and excessive notifications.
Connects to the budget module (3.5).
Custom frequency; review active crons.
💰 Budget & Tokens · 73% Is Fixed Overhead
The dirty secret of agents: most of every request is fixed cost. Knowing this changes how you operate.
About 73% of each request is fixed overhead (system prompt, tools, memory) — only about 27% is your question.
Explains why short questions still cost a lot.
High base cost; every call pays the “toll.”
The pocket guide: ~10 tokens equal about 7 words (≈70-75%).
Gives you a sense of the size before you send a huge block of text.
A token ≈ a piece of a word; it’s easy to estimate.
Clear the session often, use the right model, don’t accumulate useless skills, and keep system prompts short.
Small habits cut the bill in half.
One conversation, one goal; always start fresh.
Someone burned through 4 million tokens in 2 hours of light use; someone else spent 21,000 just asking about the weather because of a bug.
With an API key, money disappears fast if you don't keep an eye on it.
Uncontrolled loops are expensive; monitor usage.
Be specific about the model: heavy reasoning on the expensive one, volume on the cheap or free one.
The wrong choice multiplies the cost without adding any benefit.
Expensive only where it pays off; cheap everywhere else.
Set a cap (e.g., US$10/month) and the system stops when it reaches the limit.
A cap is your safety net against billing surprises.
Hard limit; alert before it’s exceeded.
🌐 Operating System · The system for everything
Hermes isn't just a chatbot: it's a single dashboard for managing personas, memory, spending, goals, and connections.
Hermes can be the layer where you manage ALL your AI in one place — the "everything of AI".
Shifts the mindset from “chat app” to “operating system.”
A hub, not a standalone feature.
Pantheon displays your personas (each with a job and model) visually within the OS.
Viewing the personas in one place makes it easier to choose who uses what.
Personas as cards; model per persona.
A memory vault that retrieves any email, meeting, or note; you can connect NotebookLM.
Centralizes the knowledge the agent can consult.
Universal recall; pluggable sources.
At a glance: available connections, model in use, memory, AI spending, gains from skills, and real-time usage.
Operating without a dashboard is flying blind.
Unified view; live metrics.
The OS "dreams" up improvements at night and comes back with proactive suggestions the next day.
The system evolves on its own, without you asking.
Self-optimization; ideas while you sleep.
Connect the OS to GitHub for a daily backup of the entire Hermes setup; if you lose your PC, restore everything.
Your OS is too valuable to die with the hardware.
Daily snapshot; full restoration.
🔗 One Brain · Hermes + Claude Code
Connecting Hermes’s brain to Claude Code: shared memory and understanding complete the 21 concepts.
The central idea is to connect the Hermes brand/brain to Claude Code so they can share context.
Two isolated agents repeat work and contradict each other.
One mind, two tools; context bridge.
Hermes knows your projects, clients, and decisions and is everywhere; Claude Code is the precision tool for building.
Each one shines in a role; knowing this helps you avoid using the wrong tool.
Broad context vs. precision; ubiquitous vs. focused.
When connected, both can access what the other has done; what you say on one side reaches the other.
Shared memory is what makes the connection meaningful.
Shared state; no silos.
If you tell Hermes one thing and Claude Code another, they diverge without a connection.
Divergence leads to rework and inconsistent decisions.
Silo = contradiction; bridge = coherence.
This is concept #21: it ties together everything that came before into a single view.
Seeing the big picture prepares you to operate for real.
From agent to OS; everything connected.
With the 21 concepts in hand, the next step is to build your own AI operating system.
Theory without practice doesn’t become an operation.
Start small; connect; iterate.