PTENES
MODULE 2.4

🪟 The agent model: Qwen 3 Coder 64k

The model that works for conversation isn’t always right for act. An agent needs to remember many things at once — which calls for a larger context window. In this module, you’ll get Qwen 3 Coder and, with a three-line file, create a version with 64k context ready to power Hermes.

6
Topics
~25
Minutes
Practical
Level
Hands-on
Type
1

❓ Why the agent requires a different model

In module 2.3, you downloaded a “fast” model and chatted with it. To chat, it's perfect. But Hermes doesn't just want to chat — it wants to act: read files, run commands, remember what it has already done, and plan the next steps. All of that takes up the context window. That's why the agent needs a model with 64k context, rather than the default model.

Reminder: "context window" is the model’s working memory — how much text it can keep in mind at once. We covered this in Track 1 (module 1.4). 64k tokens is roughly equivalent to 25,000–30,000 words of working space.

✓ AGENT model (64k)

  • ✓The entire task history fits.
  • ✓Leaves room for tool descriptions.
  • ✓Reads large files without “forgetting” the beginning.
  • ✓Supports multiple reasoning steps.

✗ Default chat model

  • ✗A short context fills up quickly when using tools.
  • ✗"Forget" instructions from the start of the task.
  • ✗Loses track during multi-step tasks.
  • ✗Great for conversation, weak for automation.

Key concepts

Acting vs. chatting

An agent uses tools; chat only responds with text.

Context in use

History + tools consume the context window.

64k tokens

The minimum headroom the agent comfortably asks for.

Dedicated model

A model just for the agent’s work.

2

🔎 What is Qwen 3 Coder

O Qwen 3 Coder and an open model from the Qwen family, trained with a focus on code and agent tasks: read files, follow technical instructions, write and fix programs. It's exactly the kind of model an AI OS needs underneath, because most of the agent's work is "editing files and running things."

📊 Why it's the foundation of the agent

  • •Open weights: downloads once, runs locally, with no monthly fee.
  • •Coding training: good at reading, writing, and editing files.
  • •30B size: runs on machines with generous RAM (see module 2.2).
  • •Flexible foundation: you can derive a version with more context — and that’s what we’ll do.

New here? "Coder" in the name doesn’t mean it’s only useful for programming—it means it was trained for this kind of structured task. Since an agent’s work looks a lot like programming (steps, tools, files), this profile is a perfect fit.

Key concepts

Qwen 3 Coder

Open model focused on code and agents.

Open weights

The weights are public; you run them locally.

30b tag

Variant with ~30 billion parameters.

Base model

From it, we created a 64k version.

3

🧩 The num_ctx trick (the Modelfile)

Here's the key insight: you doesn't need to another download to get more context. You take the model that already exists and create a "recipe" that tells Ollama: use this model, but with the context window set to 64k. This recipe is a file called Modelfile.

New here? One Modelfile and a text file with instructions for Ollama to build a model. Think of it as a recipe: the line FROM says which model is the base; PARAMETER adjusts a behavior. Here we only change the num_ctx (number of context tokens).

🎯 Objective

Create a text file called Modelfile (without an extension) in a folder of your choice, with exactly these two lines. The first points to the base model; the second opens the window to 65536 tokens (= 64k).

File contents Modelfile:

FROM qwen3-coder:30b
PARAMETER num_ctx 65536

How to verify: the file should contain only these 2 lines, in plain text. Check with cat Modelfile (Mac/Linux) or by opening it in Notepad. 65536 = 64 × 1024; that's the number shown in the video (context_length: 65536).

Variable: <qwen3-coder:30b> and the base model—switch only if you use a different tag/model. Everything else stays the same.

Key concepts

Modelfile

The text recipe that defines a derived model.

FROM

Points to the base model (here, qwen3-coder:30b).

PARAMETER num_ctx

Determines the context window size.

65536

64k tokens (64 × 1024).

4

🏗️ Create the derived model

With the Modelfile saved, a single command builds the new model. Ollama reads the recipe, reuses the weights already on your disk (it doesn't download them again), and registers a new model called qwen3-coder-64k.

🎯 Objective

Run this command in the same folder where the Modelfile is. The -f Modelfile says which recipe to use; qwen3-coder-64k and the name your new model will have.

ollama create qwen3-coder-64k -f Modelfile

How to verify: the terminal shows progress lines and ends with success. It's fast — it doesn't download the model again, it just creates the new configuration on top of what's already there.

Variable: the name qwen3-coder-64k and choose your own—use any label that helps you remember "this is the agent's one, with 64k."

Video screen showing the Modelfile with PARAMETER num_ctx 65536 and the conversation about the qwen3-coder-64k model with 64k context
Video frame: notice the PARAMETER num_ctx 65536 and in the name of the derived model (qwen3-coder-64k). This is the version the agent will use, not the default model — because it has the 64k window open.

Key concepts

ollama create

Builds a model from a Modelfile.

-f Modelfile

Points to the recipe file to use.

Weight reuse

It doesn’t download again; it uses the model already on disk.

qwen3-coder-64k

The name of the new derived model.

5

✅ Check that it worked

Before connecting to Hermes, it’s worth confirming that the model exists and that it really has 64k. Two commands do the trick: one lists the models, and the other shows the details—including the context length.

🎯 Objective

Confirm that qwen3-coder-64k appears in the list and inspect its context.

ollama list
ollama show qwen3-coder-64k

How to verify: in ollama list the name qwen3-coder-64k should appear in the table. In ollama show look for context_length: 65536 (or the parameter num_ctx 65536). If you see 65536, it's ready.

qwen3-coder:30b base model default context Modelfile FROM qwen3-coder:30b num_ctx 65536 qwen3-coder-64k 64k window open ready for the agent ✓ the base model isn’t downloaded again—it just gets a larger context window

The recipe (Modelfile) gets the base model and opens the window to 64k, producing the qwen3-coder-64k. Notice: there’s no new download along the way — the context boost comes from configuration, not reinstallation.

Key concepts

ollama list

Table of installed models.

ollama show

Model details, including context.

context_length

It should read 65536 = 64k confirmed.

Check first

Prove 64k before connecting to Hermes.

6

⚖️ 64k costs memory

There’s no free lunch: opening the window to 64k uses more RAM. The larger the context, the more memory Ollama reserves to hold everything the model is "reading" at once. That’s why the headroom rule (module 2.2) applies here again.

⚠️ The mistake to avoid

Forcing 64k on a machine with tight RAM can slow the model down or make the system use disk as memory (swap). If it freezes, reduce the num_ctx (e.g., 32768) or use a smaller base model. Context means capability, but weights matter too.

💡 Practical tip

Leave some headroom: the agent needs the model loaded AND the full context at the same time. If your machine is low-spec, test with the default model first and increase the num_ctx gradually. Downloading, testing, and deleting remain inexpensive.

Key concepts

Context uses RAM

Larger window = more memory reserved.

Headroom

Leave some RAM free for the full context.

Swap = slow

Without RAM, the system uses disk and slows to a crawl.

Adjust num_ctx

Reduce to 32768 if you run out of memory.

Optional self-check: Why did we create qwen3-coder-64k instead of using the default chat model?

🎯 Module summary

✓
The agent asks for 64k — history + tools take up the window; the chat model is too small.
✓
Qwen 3 Coder is the foundation — open model focused on code/agents, great as the engine behind Hermes.
✓
2-line Modelfile — FROM + PARAMETER num_ctx 65536, then ollama create qwen3-coder-64k -f Modelfile.
✓
Check and weigh — ollama show displays 65536; remember that 64k uses more RAM.

Next module:

2.5 — Connect the local model to Agent Hermes