🌱 Fundamentals
Start from scratch, without any jargon: what exactly is a "Jarvis," what is the LLM that thinks behind it, and the exact moment when AI stops just answering and starts ACTING for you. Three modules that unlock everything else in the course.
Read from left to right—the trail map: you start from chatbot that only responds (1.1), understands the LLM like the brain and the context as your working memory (1.2), and reaches agent that thinks and ACTS using tools real (1.3).
Learning path map
Detailed content
🤖 What is a “Jarvis” — from fiction to your home
The Iron Man assistant fantasy has become something you can build. Here, you’ll understand the concept without jargon: why a chatbot isn’t Jarvis, what an “AI operating system” is, and what you can (and still can’t) do today.
Iron Man’s Jarvis is the fantasy of an assistant that talks, remembers everything, and DOES tasks. In this course, “Jarvis” = a personal AI assistant that you build yourself — and today this is real, not fiction.
Having the right picture in your head avoids disappointment and overpromising: you’ll aim for what’s possible to build now.
Personal AI assistant: chat, remember, and act, from fantasy to something you can build.
A chatbot like ChatGPT responds. A Jarvis acts in the world: it sends an email, reads your calendar, runs a task on its own. The difference isn't how smart the text is — it's whether it makes something happen.
It's the distinction that organizes the entire course; without it, everything else becomes just "a better chat."
Responding vs. acting, taking action in the world, "the leap" that defines a Jarvis.
Just as Windows organizes programs, memory, and peripherals, a AI operating system organizes the model, memory, tools, and channels so the assistant can "live." It's the home that brings all the pieces together.
This is exactly what you’ll build in the following tracks; having the metaphor ready makes everything easier.
AI ONLY, model + memory + tools + channels, orchestration.
Three things came together: good-enough models, capable personal hardware, and open standards (such as the MCP, which lets tools connect to any agent). That’s why you can build today what was fiction yesterday.
Getting in early on a wave that's still forming is where the advantage lies—you learn before it becomes obvious.
Capable models, personal hardware, open standards (MCP), timing.
CAN do: text, search, calendar, drafts, automations. CAN'T (yet) do: long tasks without supervision—and it hallucinates (makes things up confidently). Calibrating this gets you halfway there.
Knowing the limits helps you avoid the two classic mistakes: trusting too much or giving up too soon.
Serious cases, hallucination, human supervision, realistic expectations.
Fundamentals (here) → Overview (what already exists) → Anatomy (the 6 layers) → Build → Mobile → Children. Each track builds on the previous one; the order was carefully planned.
Seeing the whole path gives you context so each part makes sense at the right time.
6 tracks, progression T1→T6, from concept to building.
🧠 The brain of the thing: what an LLM is (without jargon)
The engine that thinks behind every Jarvis, demystified: what an LLM is, how it “thinks” by predicting the next word, what context and tokens are, and the choice between running it in the cloud or on your computer.
A LLM ("Large Language Model") is a program trained on A LOT of text that learned to predict the next word. It’s what runs behind ChatGPT and Claude — the brain that "thinks in text".
It's the central piece of any Jarvis; understanding what it is demystifies everything else.
LLM, next-word prediction, trained on text, “brain.”
It doesn't query a database or “search”: it statistical forecast learned during training. That's why it's creative AND sometimes gets things wrong with complete confidence — what we call hallucination.
Knowing it's a prediction explains both the magic and the mistakes — and teaches you to check answers.
Statistical prediction, not a database, hallucination.
A context window and it’s everything the model "is seeing right now": your conversation + instructions + files. It’s like the RAM of the computer — limited, and it clears when you close the conversation.
Explains why it "forgets" — and lays the groundwork for the persistent memory in Track 3.
Context window, working memory, RAM analogy, forgotten when closed.
The model reads and writes to tokens — word pieces (a common word usually takes about one token; long words become several). The cloud charges by token, and the context window is measured in tokens.
And it’s the unit used to measure cost and context; it shows up in every practical decision.
Tokens, tokenization, token ≠ word, cost per token.
Cloud (Claude, GPT): powerful, you pay per use. Local (via Ollama): free after downloading, private, but requires hardware. Both serve the same purpose: being the brain.
It’s the trade-off that will come up throughout the course—privacy/cost on one side, power/convenience on the other.
Cloud vs. local, Ollama, pay-per-use vs. hardware, privacy.
On its own, the LLM hallucinates, doesn’t know things that happened after its training, and forgets what falls outside the context window. The solution: give it tools (search, read) and memory — exactly what Track 3 teaches.
Every limitation has a concrete solution; knowing what it is helps you move past frustration.
LLM limits, training cutoff, tools + memory as the solution.
⚙️ From chatbot to agent: when AI starts taking action
The course’s defining moment: when the LLM stops just responding and starts to ACT. Here you’ll understand the agentic loop, what tools are, and the “LLM as an operating system” metaphor driving the entire industry.
When you give an LLM tools and let it use them in a loop, it becomes an agent: an LLM that can ACT, not just respond. That’s the leap that turns chat into an assistant.
It’s at the center of everything: it’s the definition of "agent" that carries the course through to the end.
Agent, LLM + tools in a loop, answering vs. doing.
O agentic loop and the cycle: receives the request → thinks → calls a tool → reads the result → calls more tools if needed → responds. There's a iteration limit for safety, so it doesn't run forever.
It’s the mechanism that makes an agent “work on its own”; understanding it demystifies autonomy.
Agentic loop, think-act-read-repeat, iteration limit.
Tools (or “tools”) are the agent’s hands and eyes: searching the web, running code, reading a file, sending a message. Without tools, the model “responds like a stranger” — relying only on what it knows by heart.
They connect the brain to the real world; all of Track 3 revolves around them.
Tool, hands and eyes, action in the world, "reply as a stranger".
Andrej Karpathy (2023) proposed: the LLM is the kernel (core) of a new operating system. Context = RAM, tools = peripherals, agents = processes. The industry adopted this mental model.
It's the metaphor that connects "AI OS" to what you already know about computers—and underpins the course.
LLM-OS, kernel, context=RAM, tools=peripherals, agents=processes.
The idea became reality: AIOS (a real “kernel” for agents), OpenAI Operator (an agent that navigates screens) and Google Project Jarvis (an agent that operates Chrome for you). Proof that the mental model works.
Shows that this isn't isolated hype — major labs are building in the same direction as you.
AIOS, Operator, Project Jarvis, from metaphor to product.
In practice, you stop "asking" and start "delegating": instead of requesting a text, you hand over a task. It's the direct bridge to Anatomy (Track 3), where you assemble the 6 layers that make this possible.
Wraps up the Fundamentals with the change in mindset that unlocks the rest of the course.
Asking vs. delegating, a shift in approach, a bridge to the Anatomy.