🧩 Framework vs. brain
Hermes is the agentic framework (the structure that acts); the "brain" is a pluggable modelThe command /model switches this brain without changing the rest.
Switching the brain (illustrative)
/model opus-4.8 ← raciocínio pesado /model gpt ← volume geral /model deepseek ← quase de graça
🧬 Model-agnostic
Since the framework doesn’t depend on a specific model, you’re free to use the best brain for each task. This separation is the heart of the multi-brain strategy.
🧠 For reasoning: Opus with a cap
For heavy reasoning, use Opus 4.7/4.8 via OpenRouter, with spending limit (e.g., US$10/month and stop). Use an expensive model only where it’s worth it.
📊 How to control costs
- •Set a monthly cap (e.g., US$10) — when reached, it stops
- •Save Opus for difficult tasks, not everything
- •Access it via OpenRouter to compare costs easily
💡 Practical tip
The spending limit is your safety net. Without it, heavy reasoning can burn through money fast (see Track 3, the budget module).
⚡ For volume: GPT and Grok
For high-volume general tasks: use GPT (your ChatGPT via OAuth uses the US$20 subscription) or Grok (with X connected, it searches Twitter).
GPT via OAuth
Connect your ChatGPT and make use of the US$20 subscription you already pay for.
Grok + X
Good for volume and, with X connected, direct search on Twitter.
💡 Practical tip
Use what you already subscribe to. Connecting ChatGPT via OAuth (module 1.5) turns your existing subscription into Hermes's "high-volume brain."
🆓 Almost free: DeepSeek and free
DeepSeek and free models run for almost nothing. The key data point: the DeepSeek V4 flash delivers ~95% of the performance for ~1% of the cost.
📊 The number that changes everything
- ~95% of a frontier model’s performance
- ~1% of the cost
- Ideal for autopilot and high-volume background tasks
✓ Use a cheap model for
- ✓Background tasks, at scale
- ✓Autopilot and repetitive automations
- ✓When 95% is good enough
✗ Avoid in
- ✗High-impact reasoning
- ✗Critical decisions where 5% matters
- ✗Tasks that require the best brain
🔌 OpenRouter: one hub, many models
O OpenRouter is 1 connection that gives access to hundreds of models, with performance and cost rankings to compare before choosing.
🗺️ Why centralize
- •One key/connection for many models
- •Rankings help you find the best value
- •Switch models (
/model) becomes trivial
Where to Compare
# rankings de modelos (performance × custo) openrouter.ai/models
🔨 "To a hammer, everything looks like a nail"
Using a single model for everything is like having only one hammer: everything becomes a nail. The multi-brain strategy is to vary the tool according to the task.
Difficult reasoning → Opus
With a spending limit, on what really matters.
General volume → GPT/Grok
Making use of subscriptions you already have.
Background tasks → DeepSeek/free
~95% of the performance for ~1% of the cost.
💡 Practical tip
Think of Hermes as a toolbox, not a hammer. Switching brains for each task is what gets you the best result at the lowest cost.
🧯 Common mistakes when choosing a model
Most cost waste comes from using the wrong brain for the task. See what to do and what to avoid.
✓ Do it
- ✓Expensive model only for high-impact tasks
- ✓Set a spending cap on Opus
- ✓Check rankings before choosing
✗ Avoid
- ✗Use Opus for everything (burns money)
- ✗Ignore subscriptions you already pay for
- ✗Stick to a single model (the "hammer")
Pocket guide (illustrative)
raciocínio difícil -> Opus 4.7/4.8 (teto US$10/mês) volume geral -> GPT (OAuth) / Grok (com X) tarefa de fundo -> DeepSeek V4 flash (~95% perf, ~1% custo) acesso a tudo -> OpenRouter (1 conexão, rankings)
📌 Module Summary
Next Module:
1.7 — Local & Private