🪟 Context window and parameters
When choosing a model, you’ll compare TWO numbers: the parameters (the "B" from module 1.3) and the context window. They measure different things and shouldn't be confused. This module separates the two, explains why the agent requires a 64k window, and how that relates to your computer's memory.
🪟 What a context window is
A context window and the model's working memory: everything it can "keep in mind" at once to generate its next response. It includes your question, the instructions, the conversation history, and any material you pasted. If something doesn't fit in the window, the model simply can't see it.
New here? Think of a desk. Everything ON the desk is something the model can look at all at once; anything that doesn’t fit stays on the floor, out of sight. The context window and the size of that desk. Don’t confuse it with the persistent memory from module 1.2 (what the SO REMEMBERS between conversations): the window is only what fits RIGHT NOW, in this response.
The green box is the window: instructions, history, question, and what you pasted share the same space. Anything that overflows (the gray box on the right) is invisible to the model — that’s why the window SIZE matters.
Key concepts
Working memory: what fits in the current response.
The model sees what’s at the top; the rest stays out of view.
Instructions, history, question, and pasted material.
The window is the “now”; memory is what it remembers later.
🔤 What is a token
The window is measured in tokens, not words or letters tokens. A token is the piece of text the model processes (remember module 1.2? It’s what the model predicts each time). For Portuguese/English, here’s a rule of thumb: 1 token ≈ 3/4 of a word.
New here? Token and the text's "currency" for the model. It can be a whole word ("house"), part of a word ("un" + "happy" + "ly"), or even a punctuation mark. Since everything is counted in tokens, the window's size is measured in tokens too.
📊 Turning tokens into something concrete
- •Rule of thumb: 1 token ≈ 3/4 of a word (≈ 4 characters).
- •1,000 tokens ≈ 750 words ≈ 1.5 pages of text.
- •64k (65,536) tokens ≈ 25,000 to 30,000 words — a little book.
That “64k” is exactly the number we’re aiming for with the agent. Before explaining why, here’s the intuition: a 64k context window gives the model room to read dozens of pages at once — far more than a short conversation needs.
Key concepts
The small piece of text the model handles and counts.
The rule of thumb for estimating tokens.
≈ 25–30 thousand words; the agent’s target.
The window is measured in tokens, not words.
📏 Why the agent needs 64k
For simple chat, a small window is enough. But an agent is a different story. Remember module 1.2: agent = LLM + tools. Each piece of that TAKES UP context, and everything has to fit in the same window at the same time.
System instructions
How the agent should act, its rules and personas — this stays in the window the whole time.
Tool descriptions
Each tool (search, run, edit) has instructions that also use tokens.
Tool results
What each tool returns (a file's contents, a search result) goes back to the window.
Step history
The agent chains together several steps; it needs to remember what it has already done, and that adds up.
🧠 Small window = “amnesiac” agent
If the window is short, the agent forgets the first steps along the way, loses track of the tools’ context, and gets stuck on multi-step tasks. That’s why the course uses 64k: that's the breathing room the agent needs to really work without forgetting where it was.
Key concepts
Instructions + tools + history share the space.
Descriptions and results take up context.
Space for the agent to remember the steps.
Too little context = a lost agent on long tasks.
🔢 Parameters vs. context: two different numbers
Here’s the most common beginner confusion, and the heart of this module: parameters e context window are numbers independent. They don’t depend on each other. You need to learn to read BOTH when looking at a model.
🔢 Parameters (the "B")
- •Measures the SIZE of the brain (how much it "knows").
- •Fixed in the model: 32B is always 32B.
- •Determines capability and memory footprint.
🪟 Context window
- •Measures HOW MUCH it reads at once (the "desk").
- •It can be adjusted (we'll stretch it to 64k in T2).
- •Determines how much information fits in the “now.”
Analogy: imagine someone reading. The parameters are its intelligence and knowledge base (what it already knows); the context window and how many pages it can keep open in front of it at the same time. A very intelligent person with few pages open loses the thread; someone with a full desk but little background reads everything and understands little. The agent needs a strong background AND a big desk.
Key concepts
The two numbers are independent of each other.
How much the model “knows”; fixed.
How much it reads at once; adjustable.
When choosing, check both the “B” and the window.
⚖️ Model size vs. RAM (you need headroom)
All of this comes down to your hardware. For a model to run, it has to fit in memory — and it’s not enough to just barely fit: you need some headroom. Both the parameters and the context window use RAM, and they add up.
New here? RAM (the computer's working memory) is where the model lives while it runs. On Macs with Apple chips, it's "unified" and shared with the rest of the system. Simple rule: the model + the context window + what the system is already using have to fit in your RAM, with room to spare — otherwise, it freezes or gets extremely slow.
✓ With plenty of RAM headroom
- ✓The model loads and responds smoothly.
- ✓Leaves room for the system and other apps.
- ✓You can open a larger context window.
✗ No headroom (at the limit)
- ✗The model freezes or won’t even load.
- ✗It gets extremely slow (the system switches to disk).
- ✗A large context makes everything worse because it also uses more resources.
📊 Context also matters
Caution: stretching the window to 64k isn’t free. The larger the context, the more RAM it reserves while running. That’s why the practical choice (in Trail 2) is always a balance: parameters that fit + the window you actually need, leaving some headroom for the system.
Key concepts
Working memory where the model runs.
Just fitting isn’t enough; extra room prevents freezing.
A larger window uses more RAM while running.
Both count toward memory usage.
🎯 Choose the context for each task
The practical takeaway: you don’t need just ONE model. You can have a fast model of modest context for everyday conversations, and a agent model with a 64k window for when the work requires long-term memory and tools. Each task calls for a context adjustment.
🏃 Fast model (chat)
- •A smaller window is enough: short questions, direct answers.
- •Light on RAM, responds quickly.
- •Ideal for casual use in Track 2.
🤖 Agent model (64k)
- •Large window: instructions + tools + history.
- •It weighs more, but can handle multi-step tasks.
- •And that’s what we’ll set up in module 2.4.
Trail 2 spoiler: in module 2.4, you’ll create exactly this "agent model"—taking a coding model and telling it to use 64k of context through a small configuration file (Modelfile). For now, you just need to understand WHY: the agent needs a big desk.
Key concepts
Adjust the window to what the task requires.
Modest window for everyday use, light and fast.
64k window for long tasks with tools.
You can have several and use the right one for each task.
Optional self-check: Which statement about parameters and context window is correct?
🎯 Module summary
Next module:
1.5 — The trade-off: privacy, performance, and price