⚙️ From chatbot to agent: when AI starts taking action
A chatbot answers. An agent does. In this module, you'll understand the most important difference in the entire course: the moment when AI stops being a little box that returns text and starts using tools in a cycle to change the real world—send an email, read your calendar, finish a task on its own.
🚀 The leap: from answering to doing
In modules 1.1 and 1.2, you saw the brain: the LLM, that program trained on lots of text that predicts the next word. On its own, it's brilliant and completely stuck: it only knows how to do one thing — produce text. You ask, it answers. Like a brilliant person locked in a white room, with no window, no phone, and no hands. They can have a great conversation, but can't interact with anything outside.
This module’s leap—and the boundary between a chatbot and it’s a Jarvis — and give that person a phone, a keyboard, and a port. Suddenly, they don’t just talk about scheduling a meeting: they open your calendar and schedule. This "doing things in the world" has a name: agent.
🧭 The definition you’ll take with you throughout the course
A agent isn't a "smarter" model. It's the same An LLM, gaining two things:
- •Tools — ways to interact with the real world (search the web, read a file, send a message).
- •A cycle — the freedom to use these tools in sequence, deciding the next step each round, until the task is done.
In one sentence: agent = LLM + tools + a loop that decides what to do next.
✗ Chatbot — only RESPONDS
- ✗"Here's an email template you can send..."
- ✗Returns text and stops. The real work is left to you.
- ✗It can't see your calendar or access your files.
- ✗Every action depends on you copying, pasting, and clicking.
✓ Agent — ACTS in the world
- ✓“Done, I sent the email to Ana and scheduled the 3 p.m. call.”
- ✓Executes the task from start to finish and lets you know the result.
- ✓Read your calendar, search the web, open files—with permission.
- ✓You delegate; it handles the steps on its own.
New here? Chatbot = an AI program that only exchanges messages with you (ChatGPT on screen is the classic example). Agent = an AI that, in addition to chatting, can execute actions using tools in a loop until a task is complete. The difference isn’t “being smarter”—it’s having hands e a freedom to use them in sequence.
Key concepts
AI that only chats: you ask, it answers, that’s it.
LLM + tools + a loop: it does more than talk—it takes action.
The agent’s hallmark: actually changing something (email, calendar, file).
You stop asking for answers and start handing off tasks.
🔁 The Agentic Loop
An agent’s heart is a simple cycle, repeated until the task is done. We call it agentic loop. Instead of a single response, the agent runs in loops: on each loop, it thinks, perhaps uses a tool, reads what came back, and decides on the next step. When it has what it needs, it stops and responds to you.
In green, the model’s reasoning; in cyan, the part that touches the world (the tool). The dashed arrow is loop: the agent can think again as many times as needed — limited by a iteration limit so it doesn’t run forever.
📊 A concrete example: "what's the forecast for tomorrow in Salvador?"
- 1.Think: "I don't know today's forecast — I need to look it up."
- 2.Calls the tool of web search for “weather forecast Salvador tomorrow.”
- 3.Read the result: "27°C, scattered showers in the afternoon".
- 4.Decides: "I have what I need." It stops the loop and responds with the forecast.
If something were missing (e.g., confirming the city), it would take another turn. This back-and-forth is what sets an agent apart from a one-off response.
⚠️ Why there’s an iteration limit
Without a limit, a confused agent can get stuck in an infinite loop — searching, rereading, searching again — and in the cloud, spending money with every cycle. That’s why every serious system sets a ceiling (e.g., “no more than 10 rounds”). It’s a seat belt: better for it to stop and say “I couldn’t do it” than to run forever. We’ll return to this in Anatomy (Trail 3) and in “Operate” (Trail 4).
Key concepts
Think → call a tool → read → decide → repeat or respond.
One cycle iteration. The agent can do several.
Turn limit to prevent infinite loops and uncontrolled spending.
"I have what I need" — the agent decides when it's done.
🤲 Tools = the hands
If the loop is the cycle, the tools (in English, tools) are what give each loop its power. A tool is simply a action the agent can request: search the web, run a piece of code, read a file, send a message, check your calendar. They’re Jarvis’s hands and eyes.
There’s one phrase that sums up why this matters: without tools, the model "answer like a stranger". It knows nothing about you—it hasn’t seen your calendar, read your emails, or learned about your files. It gives generic advice. With tools, it accesses the your real data and becomes a real assistant.
🔧 How the agent "asks" for a tool (tool calling)
The model doesn’t execute anything on its own — it doesn’t know how to open a browser. What it does is ask: generates structured text saying "I want to use the tool buscar_web with the term Salvador forecast". The system around it (your Jarvis) is what actually executes the task, gets the result, and returns it to the model on the next loop. This mechanism has a name: tool calling (or function calling — “function call”).
Think of the model as a boss who gives orders and an assistant (the system) who goes out and carries them out. The boss never touches the tools—they decide which to use is when.
Examples of common tools
Search the web
Solves the "it doesn't know recent things" problem: brings in news, prices, today's forecast.
Read / write files
Open one of your documents, summarize it, save a note. Jarvis starts working with YOUR material.
Run code
Calculate something exactly, generate a chart, automate something. (Powerful capability — requires care; we’ll cover it in T4.)
Send a message / schedule
Send an email, schedule an event, notify you on Telegram. Here, AI actually ACTS in the world.
New here? Tool = an isolated action that the agent can trigger. Tool calling (or function calling) = the mechanism by which the model “requests” a tool instead of executing it directly. Later on (Track 3), you’ll see the MCP, a standard that securely connects tools to any agent — the “USB for AI tools.”
Key concepts
An action the agent can request: search, read, send, run.
The model REQUESTS the action; the system EXECUTES it and returns the result.
Without tools, the model only gives generic advice — it can’t see your data.
The metaphor: tools give a body to the brain that only spoke.
🧠 The LLM-OS metaphor (Karpathy)
In September 2023, Andrej Karpathy — one of the most influential AI researchers, formerly of OpenAI and Tesla — proposed an image that changed how the entire industry thinks about the subject. The idea: an LLM isn’t "just a chatbot." It’s the kernel for a new operating system. Karpathy called this LLM-OS.
New here? O kernel and it’s the core of an operating system (such as Windows or Android)—the central part that coordinates everything: memory, programs, peripherals (printer, screen, keyboard). Saying that "the LLM is the kernel of an AI OS" means that it is the coordinating heart of a new system, with the agent in control of the pieces around it.
The metaphor’s strength lies in the correspondences. Every part of an ordinary computer has an equivalent in the agent’s world—and you’ve already encountered all of them in previous modules and this one:
At the center, in green, the LLM as kernel (the coordinator). In cyan, the pieces it controls: the context window acts as RAM (the working memory from module 1.2), the tools are the peripherals, and the agents are the processes. Same design as always — just swapping the names.
🗺️ The translation table
- •LLM → o kernel (the core that coordinates everything).
- •Context window → a RAM (working memory, limited, clears when you close it—you saw this in module 1.2).
- •Tools → them peripherals (keyboard, screen, printer — whatever touches the real world).
- •Agents / subtasks → them processes (jobs that run and are coordinated).
Why does this matter to you as a non-expert? Because this metaphor became the mental model of the entire industry. When someone talks about “Agentic OS” or “AI Operating System” — including this course title — they’re using Karpathy’s idea. You don’t need to build anything now; just keep the image in mind: the LLM coordinates memory, tools, and tasks, like an OS coordinates a computer.
Key concepts
Karpathy’s metaphor (2023): the LLM as the kernel of a new OS.
The coordinating core of an operating system.
The model’s working memory, equivalent to RAM.
The "Agentic OS" that the whole industry (and this course) adopted.
🌐 AIOS, Operator, Project Jarvis
Karpathy’s metaphor didn’t stay on paper. In just a few years, it became a real product and research project. It’s worth knowing three examples—not to use now, but to see that an "agent" and an "AI OS" are the real direction the largest companies and labs are heading.
AIOS—the kernel became a paper
COLM 2025Researchers built a real kernel for AI agents: a system that organizes several agents running at the same time, with a scheduler (who runs when), memory, and storage—exactly the components of an OS. Result: up to 2.1× faster. The “LLM-OS” metaphor stopped being just an image and became a measured system.
OpenAI Operator
OpenAIAn agent that uses the computer like a human: it sees the screen, moves the cursor, clicks, and types to perform tasks in the browser (make reservations, shop, fill out forms). It’s the agentic loop with a powerful tool: the computer screen itself.
Google Project Jarvis
GoogleYes, “Jarvis” itself. A Google agent that navigates the Chrome browser for you — reading the page structure and performing steps (searching, comparing, filling things out). The name is no coincidence: the movie fantasy became a product line.
🧩 The thread that connects all three
Notice: AIOS, Operator, and Project Jarvis are the same idea as this module at different scales. They’re all agents—LLM + tools + loop. AIOS organizes many agents (the “operating system” side); Operator and Project Jarvis give an agent the broadest tool possible (the entire computer). You won’t need either one to have your own Jarvis: throughout the course, we’ll build a version lean, yours, and secure. But it’s good to know you’re on the same path as the giants.
New here? AIOS = “AI Operating System,” a research system that serves as a kernel for several agents. Operator (OpenAI) is Project Jarvis (Google) are agents that control the computer/browser by "looking" at the screen. Paper = scientific paper; COLM and an academic conference on language models. You don't need to memorize that—just know that the idea of agents is at the center of the industry.
Key concepts
Research kernel for coordinating multiple agents (2.1× faster).
OpenAI agent that sees the screen and uses the computer like a human.
Google agent that browses Chrome for you.
They’re all LLM + tools + loop — what you just learned.
🪄 What changes for you
All the theory in this module leads to a practical, very concrete change in your day-to-day life: you stop ask and starts to delegate. It’s a shift in approach—from “help me think this through” to “take care of this for me and let me know.”
✗ Before: you ASK
- ✗"Give me some subject line ideas for a payment reminder email."
- ✗You still write, open Gmail, copy, paste, and send.
- ✗AI is an advisor; the hard work is still yours.
✓ Afterward: you DELEGATE
- ✓"Charge customer X for the overdue invoice and confirm it with me."
- ✓It looks up the data, writes, sends, and gives you the result.
- ✓AI is an assistant; you supervise, you don’t execute.
📊 Honest expectations (same as module 1.1)
Delegation is powerful, but it isn't magic. An agent can still to make mistakes (remember the hallucination from module 1.2) and stumble on long tasks without supervision. That’s why a good Jarvis has brakes: confirmation for serious actions (sending money, deleting a file), an iteration limit, and a record of what it did. The point isn't to trust blindly — it's delegate with safeguards.
These safeguards are covered in Track 4 ("Build")—here, you just need to know they exist and are part of a well-built agent.
🌉 The bridge to Anatomy (Track 3)
You already have the three fundamental pieces: brain (the LLM, module 1.2), hands (the tools) and movement (the loop). That completes Track 1. From here, the course will break this agent down into six layers — channels, identity, tools, skills, agents, and brains — one at a time. It’s how you go from the idea “a Jarvis acts” to the detail of how each part works and fits together.
Self-check (optional): Which sentence best sums up what makes an AI an "agent"?
Key concepts
Complete the entire task, not just ask for an answer.
You check the result; the agent handles the steps.
Confirmation, an iteration limit, and a record of what was done.
Brain + hands + loop, opened up in 6 layers in Track 3.
🎯 Module summary
Next:
Return to Track 1—you’ve completed the Fundamentals. Next are Track 2 (The Big Picture) and Track 3 (Anatomy), where the agent is opened up layer by layer.