PTENES
MODULE 3.5

💰 Budget & Tokens · 73% Is Fixed Overhead

The Dirty Secret of Agents: ~73% of each request is fixed overhead — system prompt, tools, memory. Only ~27% is your question. Understanding this changes how you operate and helps you avoid burning through 4 million tokens without realizing it.

6
Topics
~25
Minutes
Advanced
Level
Practice
Type
1

🥧 73% fixed overhead

When you send a one-line question, the model doesn't read only that. It reads all the overhead first: the system prompt, the tool list, the memory, the loaded skills. This fixed part is about 73% for each request — and you pay for it on every call, no matter how long the question is.

73% fixed overhead 27% useful Fixed overhead: system prompt, tools, memory, skills Useful: your real question → Every call pays the 73% toll

Illustrative diagram · approximate ratio of overhead × useful content

2

🔤 Tokens ↔ words

To get a sense of the cost, you need to convert it. The back-of-the-envelope calculation: ~10 tokens equal about 7 words (≈70-75%). A token is a piece of a word — common words are 1 token, while long words become 2 or 3.

Fast conversion (illustrative)

10 tokens ≈ 7 words
100 tokens ≈ ~70 words (a short paragraph)
1,000 tokens ≈ ~700 words (one page)

📊 Why estimation matters

Before you paste a 50-page document, you already know: that's ~35,000 input tokens alone — and that goes into all subsequent calls if you don’t clear the session.

3

🧹 Cost-saving strategies

The good news: small habits cut the bill in half. The secret is to keep overhead lean and the session clean.

✓ What to DO

  • ✓Clear the session often (one conversation, one goal).
  • ✓Use the right model for each job.
  • ✓Keep system prompts short.
  • ✓Compress when the context grows (from large to small).

✗ What to AVOID

  • ✗Loading dozens of skills you don't use.
  • ✗Huge system prompts "just in case".
  • ✗Keep one endless session, accumulating context.
  • ✗Use an expensive model for a trivial task.

💡 Practical tip

The material’s motto: one conversation, one goal; always clear it and start over. Every useless skill loaded as overhead is paid for in every message of the session.

4

🔥 4 million tokens in 2h

The material's real warning: someone burned 4 million tokens in 2 hours of apparently light usage. Another spent 21,000 tokens just asking the time because of an error in a loop. With an API key, money disappears fast if you don’t keep an eye on it.

!

4M tokens in 2h of light use

Context that grows without cleanup + crons firing + fixed overhead = a silent explosion in consumption.

!

21,000 tokens just to know the time

An error put the agent in a retry loop — each retry paying the 73% overhead again.

🚨 Caution

Uncontrolled loops are the biggest villain on your bill. Monitor usage and use /stop when something seems stuck repeating the same thing.

5

🎛️ The right model, the right cost

The model is the biggest cost lever. Be specific: heavy reasoning goes to the expensive model (Opus); volume and simple tasks go to the cheap or free model (DeepSeek, GPT via OAuth). Using Opus for everything multiplies the cost without improving quality.

Reasoning

Expensive model only where it’s worthwhile.

Volume

Low-cost model or subscription-based.

Routine / autopilot

Free model, nearly free.

6

📊 Spending Limits

The ultimate safety net: define a spending limitOn OpenRouter, for example, you set a limit (e.g., US$10/month), and the system simply stops when it reaches it. That lets you sleep easy even with crons running.

Usage dashboard · illustrative recreation, not a real screenshot

Monthly spendingUS$ 7.40 / 10.00
Alert in80%
73% fixed

Every paid call.

Clear session

One conversation, one goal.

Right model

Expensive only where it pays off.

Cap

Stop at the limit.

7

🧾 Where the money leaks (checklist)

Putting it all together: the biggest token drains are predictable. Run through this mental checklist whenever the bill looks alarming.

1

Infinite session

Accumulated context goes into every call. Clear it and start over.

2

Expensive model for a trivial task

Using Opus for "what time is it" is wasteful. Switch with /model.

3

Forgotten crons

A frequent schedule multiplies the overhead. Review the crons.

4

Retry loop

A loop error charges you again for the 73% on each attempt. Use /stop.

💡 Practical tip

Keep the usage dashboard open for the first few weeks. Watching the number rise in real time is the best teacher of token economy.

📌 Module Summary

✓
73% overhead - each request pays a fixed toll for the system prompt, tools, and memory.
✓
10 tokens ≈ 7 words - estimate before sending huge text.
✓
Clear and focus - one conversation, one goal; no useless skills.
✓
Loops = the villain - 4M tokens in 2h and 21k to check the time came from loops and bloated context.
✓
Spending cap - set a limit and the system stops on its own.

Next Module:

3.6 - 🌐 Operating System: the single dashboard where you manage personas, memory, spending, and goals.