PTENES
MODULE 1-2

🧠 The brain of the thing: what an LLM is (without jargon)

Behind every Jarvis, a brain is at work: the LLM. But what is that, exactly? In this module, you’ll demystify the “model” without any math—you’ll understand how it predicts the next word, why it has a short memory, what tokens are, and the difference between running in the cloud or on your machine. By the end, you’ll know exactly what it can (and can’t) do.

6
Topics
~35
Minutes
Basic
Level
Theory
Type
1

📖 What is an LLM

When you talk to ChatGPT, Claude, or Gemini, you’re talking to a LLM. The acronym comes from the English Large Language Model — in Portuguese, “Large Language Model.” Breaking it down word by word: it’s a program (model) that handles text (language) and was trained on a gigantic amount of material (large). This program is the brain of your Jarvis — the part that thinks, writes, and responds.

The central idea is almost ridiculously simple: the LLM was trained to predict the next word. You type "The sky is..." and it calculates which word is likely to come next—"blue," "clear," "cloudy." But it does this millions of times per second, word after word, and the result looks like reasoning, conversation, even creativity. It learned these patterns by reading practically the entire internet, books, and code.

🧠 The LLM in one sentence

An LLM is a giant, very well-trained autocomplete. The same feature that suggests the next word on your phone’s keyboard, taken to a level where it writes entire texts, summarizes documents, and answers questions — because it learned the patterns of human language by reading a mountain of text.

  • •Large (large): trained on a huge amount of text, with billions of “internal adjustments.”
  • •Language (language): its raw material is text—text goes in, text comes out.
  • •Model (model): a program that learned patterns, not a database of memorized answers.

New here? LLM = “Large Language Model,” the type of AI program that runs behind products like ChatGPT and Claude. Model, in AI jargon, is simply the trained “brain” — the software file that receives text and returns text. Whenever you read “model” in this course, think “the AI brain.”

Key concepts

LLM

Large Language Model — the AI brain that thinks in text.

Predict the next word

The basic mechanism: it chooses, word by word, what comes next.

Training

Learned language patterns by reading a huge amount of text.

Model = brain

In Jarvis, the LLM is the piece that reasons; the others provide hands and memory.

2

🔮 How it “thinks”

Here’s the part that most confuses beginners: the LLM doesn't query a database e doesn't do a Google search before answering. It doesn't have a "fact table" where it looks up the right answer. What it does is statistical forecast: based on what you wrote, it estimates the most likely continuation, based on patterns it learned during training.

This explains the two sides of an LLM. On one hand, it is creative and flexible: because it isn't copying memorized responses, it can write a new poem, adapt a text to your tone, and combine ideas in original ways. On the other hand, it sometimes gets it wrong with complete confidence — makes up a fact, cites a book that doesn’t exist, gives the wrong date—and says it in the same confident tone it uses when it’s right. This confident mistake has a name: hallucination.

📊 Prediction, not a photographic memory

  • •It isn't search: it doesn't "look for" the answer; it generates word for word.
  • •It isn't a database: there isn't a spreadsheet of stored facts—there are learned patterns.
  • •It’s probability: "which word usually comes here?" repeated thousands of times.
  • •That’s why it varies: the same question can produce slightly different answers each time.

⚠️ Warning: hallucination

Since its goal is always to give "the most plausible continuation," sometimes the plausible continuation simply isn't true. It may make up a phone number, a scientific paper, or a quote—and present it with certainty. So, for any fact that matters (date, amount, name, law), check against a source. In the next tracks, tools and memory will greatly reduce this risk.

New here? Hallucination and it’s when an LLM responds with something that seems right and sounds convincing, but is false. It’s not a lie in the human sense—it’s a side effect of always trying to predict "what usually comes next here," even when it doesn’t really know the answer.

Key concepts

Statistical prediction

It estimates the most likely continuation; it doesn't look up stored facts.

Hallucination

Confident mistake: something that sounds right but is made up.

Creativity

Because it doesn’t memorize answers, it generates new, adaptable text.

Check what matters

For sensitive facts, verify with a trusted source.

3

🧩 Context = working memory

If the LLM doesn’t have a database, how does it “remember” what you said at the start of the conversation? The answer is context. Everything the model is “seeing right now” — your message, earlier replies in the conversation, the instructions you gave, a file you pasted — lives on a kind of workbench called context window. And that’s the only place it gets information from when answering.

The best analogy is the Computer RAM: it's short-term memory, fast, but limited. Everything in the window is clear to the model. What leaves it (because the conversation got too long, or because you closed and reopened it) simply disappears — how to wipe the board. That’s why ChatGPT “forgets” you between conversations: each new conversation starts with an empty table.

The LLM as “kernel”: the brain at the center of a small operating system LLM the brain that thinks context window = working memory (RAM) tools = peripherals (web, files) you sends text, receives text Everything that enters the window is visible to the LLM. What leaves it, it forgets.

At the center, the LLM and the brain. On top, the context window is working memory (like RAM): fast, but limited. This metaphor — LLM at the center, context as RAM, tools as peripherals — guides the entire course and returns strongly in module 1-3.

New here? Context window and the model's "field of view": everything it can take into account right now. RAM is the computer's short-term memory, which clears when powered off. The context window is exactly that for the LLM — and that's why the lasting memory (which we’ll see in Track 3) needs to be built externally, in files.

Key concepts

Context

Everything the model is "seeing now": conversation + instructions + files.

Context window

The limited space where this context fits.

Like RAM

Working memory: fast, limited, cleared when you close it.

Forgets when you leave

Each new conversation starts with an empty workspace.

4

🔤 Tokens: the model’s currency

The LLM doesn’t read “words” the way you do. It breaks text into little pieces called tokens. A token can be a whole word ("cat"), part of a word ("unbeliev" + "able"), or even just a punctuation mark. A useful rule of thumb: in Portuguese, each token is roughly equivalent to 4 characters, and 100 words come to around 130 to 150 tokens. All the model's reading and writing happens in tokens.

Why does this matter to you? For two very practical reasons. First: the context window is measured in tokens — when someone says “this model has a context of 128 thousand tokens,” they’re describing the size of the workbench. Second: in the cloud, you pay per token — both for the tokens that come in (your question + the context) and those that go out (the response). A token is literally the LLM’s currency.

How the model "sees" the sentence: My Jarvis is amazing! My Jarvis e incri vel ! “incredible” became 2 tokens 6 tokens in total — and that's what counts.

The sentence becomes 6 tokens. Notice that “incredible” splits into two (incri + vel) and even “!” is a token. In this unit, the model reads, writes, measures the context, and—in the cloud—charges you.

📊 Numbers that help put things in perspective

  • •~4 characters = 1 token (in Portuguese, more or less).
  • •100 words ≈ 130 to 150 tokens.
  • •1 book page ≈ 400 to 600 tokens.
  • •128 thousand token context ≈ an entire ~300-page book fitting on the “desk.”

💡 Practical tip

Long conversations cost more and fill the window faster—because the model reprocesses everything with each response. For independent tasks, open a new conversation makes Jarvis faster, cheaper, and less likely to get confused by old topics. “Clearing the table” is a good practice, not a waste.

Key concepts

Token

The small piece of text the model reads and writes (a word, fragment, or symbol).

Context in tokens

Window size is measured in tokens, not words.

Per-token billing

In the cloud, you pay for input and output tokens.

Rule of thumb

~4 characters per token; 100 words ≈ 130-150 tokens.

5

☁️ Local model vs. cloud

An LLM needs to run somewhere—and there are two places. In the cloud, the model lives on a company’s server (Claude from Anthropic, GPT from OpenAI, Gemini from Google), and you talk to it over the internet. In local, the model runs on your own machine — usually through a program called Ollama, which downloads and runs open models for free. Both do the same job: they serve as the brain. What changes is where they run, the cost, and privacy.

✓ Cloud (Claude, GPT, Gemini)

  • ✓The world's most powerful models, always ready.
  • ✓It doesn't require powerful hardware—it even runs on a phone.
  • ✓No installation: just open and use.
  • ✗You pay per token, and your data leaves the machine.

✓ Local (Ollama on your machine)

  • ✓Free after download: $0 per token, no meter.
  • ✓Private: nothing leaves your machine; it works offline.
  • ✓You own the brain—without depending on anyone.
  • ✗Requires hardware (RAM/GPU), and smaller models are weaker.

The choice isn’t “one is better than the other”—it’s which is best for this task. A cheap local model handles a simple summary without cost or data leakage. A powerful cloud model shines on a difficult task that requires maximum reasoning. Best of all, your Jarvis will be able to switch brains depending on the case — something that Track 3 (module "Brains") shows in detail.

New here? Ollama and a free program that installs and runs LLMs on your own machine with a simple command. It's the most common way to have a "local brain." You'll hear a lot about it in the hands-on tracks—for now, remember: Ollama = run the LLM at home for free.

Key concepts

Cloud

The model runs on the company’s server; it’s powerful, but you pay per use.

Local

The model runs on your machine; it’s free and private, but requires hardware.

Ollama

The program that downloads and runs open LLMs locally.

Change the brain

Use local or cloud models depending on the task — decide for each response.

6

🚧 Limitations and workarounds

Now that you understand how the LLM works, you can clearly see three built-in limits—and, even better, how each one is handled. The good news: none of these limits is a dealbreaker. They’re exactly why a Jarvis has tools and memory on top of the model. The LLM alone is just the brain; the rest provides what's missing.

1

It hallucinates

Limit: makes up facts confidently. Workaround: give it tools (web search, file reading) to check reality instead of guessing.

2

It doesn't know what's new

Limit: its knowledge stops at the date it was trained (it doesn't know yesterday's news). Workaround: search tools bring current information into context.

3

Forgets outside the context

Limit: what leaves the window disappears. Workaround: persistent memory in files and search—so Jarvis "remembers" between conversations.

🧩 The brain doesn’t work alone

Notice the pattern: every limitation of a "pure" LLM is addressed by adding something around of it. Tools address hallucinations and freshness; memory addresses forgetting. That’s exactly why a Jarvis is more than one LLM — it’s the LLM more the layers that make it useful in the real world.

In module 1-3, you see the leap from "chatbot that responds" to "agent that acts," and in Track 3 (Anatomy), each of these layers becomes its own chapter.

Self-check (optional): Which sentence best describes an LLM?

Key concepts

Training cutoff

The model’s knowledge stops at the date it was trained.

Tools solve

Search and reading help prevent hallucinations and outdated information.

Persistent memory

Files + search give Jarvis memory across conversations.

LLM + layers

A Jarvis is the brain plus the layers that make it useful.

🎯 Module summary

✓
What is an LLM — the AI “brain,” a program that predicts the next word after reading a mountain of text.
✓
How it thinks — statistical prediction, not search; that’s why it’s creative and why it sometimes hallucinates.
✓
Context and tokens — the context window is the "RAM" (limited, cleared when you exit), and everything is measured and billed in tokens.
✓
Local vs. cloud and limits — same brain, different places; and the limits (hallucinating, going out of date, forgetting) become tools and memory.

Next module:

Module 1-3 — From chatbot to agent: when AI starts taking action