PTENES
Skip to content
MODULE 1 - FUNDAMENTALS

🔤 Tokens and Context Window

Tokens are the basic units LLMs process—they are not words, but pieces of text. Understanding tokens is essential to knowing the limits of what you can do.

~4
Characters/Token
200K
Claude Context
1.3x
Words → Tokens (PT)
🧩

What Are Tokens?

The basic processing unit of LLMs

Tokens are how the LLM “reads” text. A word can be 1 token or multiple tokens. The model doesn't process individual characters or complete words—it processes tokens.

📝 Tokenization Examples

"Hello"

= 1 token

"Brasília"

= 2 tokens (Bras + ília)

"ChatGPT"

= 2 tokens (Chat + GPT)

"Artificial Intelligence"

= 4-5 tokens

📏 General Rule

~4 characters = 1 token in Portuguese. Spaces and punctuation count too!

🤔 Why Do Tokens Matter?

🚫

Physical Limits

If you exceed the context window, the model doesn’t work

💰

Cost

APIs charge per token—for both input and output

⚡

Performance

Very long prompts can affect quality and speed

📦

Context Window (Context Window)

The maximum number of tokens the model processes

The “context window” is the maximum number of tokens the model can process at once. Everything needs to fit within this limit—your prompt, history, and the response.

📊 Sizes by Model (2025)

C

Claude 3.5 Sonnet

Anthropic

200.000 tokens

~150,000 words

G

GPT-4 Turbo

OpenAI

128,000 tokens

~96,000 words

G

Gemini 1.5 Pro

Google

1,000,000 tokens

~750,000 words

📋 What Counts Toward the Context Window?

  • ✓ Your current prompt
  • ✓ The entire conversation history
  • ✓ The model's response
  • ✓ Examples and context provided
🧮

Calculating Tokens in Practice

Learn how to estimate whether your content will fit

✅ Example: Short Text

A 1,000-word text in Portuguese

≈ 1.300 tokens

If the limit is 4,000 tokens:

  • • Your text: 1,300 tokens
  • • Your question: ~50 tokens
  • • Expected response: ~500 tokens
  • • Total: ~1,850 tokens ✅ Fits!

❌ Example: Large Document

50,000-word document

≈ 65.000 tokens

If the limit is 16,000 tokens (GPT-3.5):

❌ DOESN'T FIT!

💡 Solutions:

  1. 1. Use a model with a larger context window (Claude, Gemini)
  2. 2. Break the Document into Smaller Parts
  3. 3. Summarize first, then analyze the summary
⭐

Practical Tips

Optimize your token usage

💡 Use tools like tiktoken (OpenAI) to count exact tokens before sending.
💡 For Portuguese, estimate 1.3x the number of words to get an approximate token count.
💡 Always leave 20-30% margin for the response—don’t use 100% of the context with your prompt.
💡 If you exceed the limit, break it into smaller parts or use progressive summarization techniques.
💡 Remember: the conversation history counts too - long conversations consume context.
⚠️

Common Errors

What to avoid with tokens

❌ Ignoring token counts

Example: Pasting a Huge Document Without Checking Whether It Fits

✓ Solution: Always estimate tokens before sending. Use counting tools.

❌ Using the entire context window

Example: 15,900-Token Prompt in a 16k Model

✓ Solution: Leave at least 20-30% for the model's response.

🚀 Next Step

Now that you understand tokens, learn to Anatomy of a Prompt to structure your instructions effectively!

Go to Anatomy