🥧 73% fixed overhead
When you send a one-line question, the model doesn't read only that. It reads all the overhead first: the system prompt, the tool list, the memory, the loaded skills. This fixed part is about 73% for each request — and you pay for it on every call, no matter how long the question is.
Illustrative diagram · approximate ratio of overhead × useful content
🔤 Tokens ↔ words
To get a sense of the cost, you need to convert it. The back-of-the-envelope calculation: ~10 tokens equal about 7 words (≈70-75%). A token is a piece of a word — common words are 1 token, while long words become 2 or 3.
Fast conversion (illustrative)
📊 Why estimation matters
Before you paste a 50-page document, you already know: that's ~35,000 input tokens alone — and that goes into all subsequent calls if you don’t clear the session.
🧹 Cost-saving strategies
The good news: small habits cut the bill in half. The secret is to keep overhead lean and the session clean.
✓ What to DO
- ✓Clear the session often (one conversation, one goal).
- ✓Use the right model for each job.
- ✓Keep system prompts short.
- ✓Compress when the context grows (from large to small).
✗ What to AVOID
- ✗Loading dozens of skills you don't use.
- ✗Huge system prompts "just in case".
- ✗Keep one endless session, accumulating context.
- ✗Use an expensive model for a trivial task.
💡 Practical tip
The material’s motto: one conversation, one goal; always clear it and start over. Every useless skill loaded as overhead is paid for in every message of the session.
🔥 4 million tokens in 2h
The material's real warning: someone burned 4 million tokens in 2 hours of apparently light usage. Another spent 21,000 tokens just asking the time because of an error in a loop. With an API key, money disappears fast if you don’t keep an eye on it.
4M tokens in 2h of light use
Context that grows without cleanup + crons firing + fixed overhead = a silent explosion in consumption.
21,000 tokens just to know the time
An error put the agent in a retry loop — each retry paying the 73% overhead again.
🚨 Caution
Uncontrolled loops are the biggest villain on your bill. Monitor usage and use /stop when something seems stuck repeating the same thing.
🎛️ The right model, the right cost
The model is the biggest cost lever. Be specific: heavy reasoning goes to the expensive model (Opus); volume and simple tasks go to the cheap or free model (DeepSeek, GPT via OAuth). Using Opus for everything multiplies the cost without improving quality.
Expensive model only where it’s worthwhile.
Low-cost model or subscription-based.
Free model, nearly free.
📊 Spending Limits
The ultimate safety net: define a spending limitOn OpenRouter, for example, you set a limit (e.g., US$10/month), and the system simply stops when it reaches it. That lets you sleep easy even with crons running.
Usage dashboard · illustrative recreation, not a real screenshot
Every paid call.
One conversation, one goal.
Expensive only where it pays off.
Stop at the limit.
🧾 Where the money leaks (checklist)
Putting it all together: the biggest token drains are predictable. Run through this mental checklist whenever the bill looks alarming.
Infinite session
Accumulated context goes into every call. Clear it and start over.
Expensive model for a trivial task
Using Opus for "what time is it" is wasteful. Switch with /model.
Forgotten crons
A frequent schedule multiplies the overhead. Review the crons.
Retry loop
A loop error charges you again for the 73% on each attempt. Use /stop.
💡 Practical tip
Keep the usage dashboard open for the first few weeks. Watching the number rise in real time is the best teacher of token economy.
📌 Module Summary
Next Module:
3.6 - 🌐 Operating System: the single dashboard where you manage personas, memory, spending, and goals.