PTENES
MODULE 1.4

🪟 Context window and parameters

When choosing a model, you’ll compare TWO numbers: the parameters (the "B" from module 1.3) and the context window. They measure different things and shouldn't be confused. This module separates the two, explains why the agent requires a 64k window, and how that relates to your computer's memory.

6
Topics
~30
Minutes
Basic
Level
Theory
Type
1

🪟 What a context window is

A context window and the model's working memory: everything it can "keep in mind" at once to generate its next response. It includes your question, the instructions, the conversation history, and any material you pasted. If something doesn't fit in the window, the model simply can't see it.

New here? Think of a desk. Everything ON the desk is something the model can look at all at once; anything that doesn’t fit stays on the floor, out of sight. The context window and the size of that desk. Don’t confuse it with the persistent memory from module 1.2 (what the SO REMEMBERS between conversations): the window is only what fits RIGHT NOW, in this response.

context window (the “desk”) instructions history your question pasted document free space the model can see everything here at once what didn’t fit (stays out of sight)

The green box is the window: instructions, history, question, and what you pasted share the same space. Anything that overflows (the gray box on the right) is invisible to the model — that’s why the window SIZE matters.

Key concepts

Context window

Working memory: what fits in the current response.

The "desk"

The model sees what’s at the top; the rest stays out of view.

What goes in

Instructions, history, question, and pasted material.

≠ OS memory

The window is the “now”; memory is what it remembers later.

2

🔤 What is a token

The window is measured in tokens, not words or letters tokens. A token is the piece of text the model processes (remember module 1.2? It’s what the model predicts each time). For Portuguese/English, here’s a rule of thumb: 1 token ≈ 3/4 of a word.

New here? Token and the text's "currency" for the model. It can be a whole word ("house"), part of a word ("un" + "happy" + "ly"), or even a punctuation mark. Since everything is counted in tokens, the window's size is measured in tokens too.

📊 Turning tokens into something concrete

  • •Rule of thumb: 1 token ≈ 3/4 of a word (≈ 4 characters).
  • •1,000 tokens ≈ 750 words ≈ 1.5 pages of text.
  • •64k (65,536) tokens ≈ 25,000 to 30,000 words — a little book.

That “64k” is exactly the number we’re aiming for with the agent. Before explaining why, here’s the intuition: a 64k context window gives the model room to read dozens of pages at once — far more than a short conversation needs.

Key concepts

Token

The small piece of text the model handles and counts.

≈ 3/4 of a word

The rule of thumb for estimating tokens.

64k tokens

≈ 25–30 thousand words; the agent’s target.

Window unit

The window is measured in tokens, not words.

3

📏 Why the agent needs 64k

For simple chat, a small window is enough. But an agent is a different story. Remember module 1.2: agent = LLM + tools. Each piece of that TAKES UP context, and everything has to fit in the same window at the same time.

🧾

System instructions

How the agent should act, its rules and personas — this stays in the window the whole time.

🛠️

Tool descriptions

Each tool (search, run, edit) has instructions that also use tokens.

📥

Tool results

What each tool returns (a file's contents, a search result) goes back to the window.

🔁

Step history

The agent chains together several steps; it needs to remember what it has already done, and that adds up.

🧠 Small window = “amnesiac” agent

If the window is short, the agent forgets the first steps along the way, loses track of the tools’ context, and gets stuck on multi-step tasks. That’s why the course uses 64k: that's the breathing room the agent needs to really work without forgetting where it was.

Key concepts

Everything in the same window

Instructions + tools + history share the space.

Tools cost tokens

Descriptions and results take up context.

64k = room to spare

Space for the agent to remember the steps.

Short window stalls

Too little context = a lost agent on long tasks.

4

🔢 Parameters vs. context: two different numbers

Here’s the most common beginner confusion, and the heart of this module: parameters e context window are numbers independent. They don’t depend on each other. You need to learn to read BOTH when looking at a model.

🔢 Parameters (the "B")

  • •Measures the SIZE of the brain (how much it "knows").
  • •Fixed in the model: 32B is always 32B.
  • •Determines capability and memory footprint.

🪟 Context window

  • •Measures HOW MUCH it reads at once (the "desk").
  • •It can be adjusted (we'll stretch it to 64k in T2).
  • •Determines how much information fits in the “now.”

Analogy: imagine someone reading. The parameters are its intelligence and knowledge base (what it already knows); the context window and how many pages it can keep open in front of it at the same time. A very intelligent person with few pages open loses the thread; someone with a full desk but little background reads everything and understands little. The agent needs a strong background AND a big desk.

Key concepts

Independent

The two numbers are independent of each other.

Parameters = repertoire

How much the model “knows”; fixed.

Context = the desk

How much it reads at once; adjustable.

Read both

When choosing, check both the “B” and the window.

5

⚖️ Model size vs. RAM (you need headroom)

All of this comes down to your hardware. For a model to run, it has to fit in memory — and it’s not enough to just barely fit: you need some headroom. Both the parameters and the context window use RAM, and they add up.

New here? RAM (the computer's working memory) is where the model lives while it runs. On Macs with Apple chips, it's "unified" and shared with the rest of the system. Simple rule: the model + the context window + what the system is already using have to fit in your RAM, with room to spare — otherwise, it freezes or gets extremely slow.

✓ With plenty of RAM headroom

  • ✓The model loads and responds smoothly.
  • ✓Leaves room for the system and other apps.
  • ✓You can open a larger context window.

✗ No headroom (at the limit)

  • ✗The model freezes or won’t even load.
  • ✗It gets extremely slow (the system switches to disk).
  • ✗A large context makes everything worse because it also uses more resources.

📊 Context also matters

Caution: stretching the window to 64k isn’t free. The larger the context, the more RAM it reserves while running. That’s why the practical choice (in Trail 2) is always a balance: parameters that fit + the window you actually need, leaving some headroom for the system.

Key concepts

RAM

Working memory where the model runs.

Needs some slack

Just fitting isn’t enough; extra room prevents freezing.

Context is heavy

A larger window uses more RAM while running.

Parameters + context window

Both count toward memory usage.

6

🎯 Choose the context for each task

The practical takeaway: you don’t need just ONE model. You can have a fast model of modest context for everyday conversations, and a agent model with a 64k window for when the work requires long-term memory and tools. Each task calls for a context adjustment.

🏃 Fast model (chat)

  • •A smaller window is enough: short questions, direct answers.
  • •Light on RAM, responds quickly.
  • •Ideal for casual use in Track 2.

🤖 Agent model (64k)

  • •Large window: instructions + tools + history.
  • •It weighs more, but can handle multi-step tasks.
  • •And that’s what we’ll set up in module 2.4.

Trail 2 spoiler: in module 2.4, you’ll create exactly this "agent model"—taking a coding model and telling it to use 64k of context through a small configuration file (Modelfile). For now, you just need to understand WHY: the agent needs a big desk.

Key concepts

Context per task

Adjust the window to what the task requires.

Fast model

Modest window for everyday use, light and fast.

Agent model

64k window for long tasks with tools.

More than one model

You can have several and use the right one for each task.

Optional self-check: Which statement about parameters and context window is correct?

🎯 Module summary

✓
Context window — working memory: what the model can see at once; measured in tokens.
✓
Token — ≈ 3/4 of a word; 64k ≈ 25–30 thousand words.
✓
64k in the agent — instructions, tools, and history share the window; you need some headroom.
✓
Parameters ≠ context + RAM — two independent numbers; both affect memory, so choose based on the task.

Next module:

1.5 — The trade-off: privacy, performance, and price