PTENES
MODULE 2-3

🌍 The world out there (AIOS, MemGPT, Operator) and the way forward

The family you saw in the previous module didn’t appear out of nowhere. Out there, labs, academic papers, and tech giants are building the SAME ideas—AI kernels, infinite memory, agents that run code, pocket-sized hardware. In this module, you’ll visit this “world out there,” understand the MCP standard that ties it all together, and wrap up with an honest recommendation on where YOU should start.

6
Topics
~35
Minutes
Beginner
Level
Overview
Type
1

📄 AIOS — when the metaphor became a paper

In Track 1, you heard Andrej Karpathy's phrase: "the LLM is the kernel of a new operating system.” It was a metaphor—a nice way to think about it. Well, in 2024 a group of researchers decided to take the metaphor seriously and BUILD that system. The result is called AIOS (for “LLM Agent Operating System,” or simply “agent operating system”), and became an academic paper accepted at COLM 2025 (one of the field’s serious conferences).

New here? A “kernel” (core) is the central part of an operating system like Windows or Android—the piece that allocates memory, decides which program runs now, and communicates with peripherals (keyboard, screen, disk). A "paper" and a published scientific paper: a way for the community to say "this was tested and reviewed; it's not just marketing talk."

The key idea behind AIOS is easy to understand: when you have MANY AI agents trying to run at the same time, they fight over the same resources—they all want the model, memory space, and somewhere to store files. A regular operating system already handles this kind of conflict for ordinary programs. AIOS does the same, but for AI agents: it puts a scheduler (which one runs first), an memory manager (which one uses the context) and a storage vault (where the data stays) between agents and the LLM.

📊 The number that matters

  • •2.1x faster: in the paper's tests, organizing agents with the AIOS "kernel" made several agents finish more than twice as fast as when running in a mess.
  • •Open source: isn't a closed product; it's a reference anyone can study.
  • •The lesson for you: the idea “LLM = kernel” stopped being poetry and became real engineering.

🧭 Why this interests you

You won’t install AIOS on your personal Jarvis—it’s heavyweight and aimed at people running dozens of agents. But it PROVES that the way this course thinks about things (model + memory + tools organized as an OS) is the same one academia is formalizing. When you build a lean Jarvis at home, you’re building a "pocket-sized" version of the same idea found in the papers.

Key concepts

AIOS

An “operating system for AI agents”: scheduler + memory + storage.

Kernel

The core of an OS; here, the LLM at the center of everything.

Scheduler

Decides which agent runs first when several are competing.

COLM 2025

Serious academic review—the seal of “this has been reviewed.”

2

🧠 MemGPT / Letta — giving the agent “infinite memory”

You already know from Track 1 that the LLM has a context window limited—a kind of working memory that fills up and "forgets" what no longer fits. This is a huge problem: how will Jarvis remember you three months from now if it forgets everything when the conversation gets long? The project MemGPT (now renamed Letta) was created specifically to solve this.

New here? Hierarchical memory means memory in layers, from the fastest and smallest to the slowest and largest — just like your computer, which has RAM (fast and expensive) and a disk (slow and cheap). MemGPT's idea is to copy this exact operating system trick and apply it to the AI agent.

Their inspiration was the memory pagination of operating systems: when RAM fills up, the OS sends the least-used chunks to disk and brings them back when needed. MemGPT makes the agent do this ON ITS OWN with its own memories: what is urgent stays "on the desk," what is old goes in the archive, and it learns to retrieve it when the subject comes up again. In practice, it divides memory into three levels:

1

Core (nucleus)

The essentials that are ALWAYS in front of the agent: who you are, what it is. Like the hottest RAM.

2

Recall (recent)

The recent conversation that still matters, but can leave the scene if it isn’t used.

3

Archival

The vault for everything that's ever been said, stored outside the context and retrieved on demand. The "infinite memory."

📊 What this means out there

The idea was so well received that the company behind it (Letta) raised around US$ 10 million investment. The market bet that “giving an agent real memory” is a valuable problem to solve.

Remember this concept: in Track 3, the “Brains” module, you'll see a SIMPLE, homegrown version of the same idea — memory in text files + a search index. You don't need US$ 10 million; you need an organized folder.

Key concepts

MemGPT / Letta

The project that gives the agent “infinite memory.”

Hierarchical memory

Layered memory: core, recall, archival.

Pagination

The operating system trick of swapping pieces between fast and slow memory.

Infinite memory

The agent retrieves old memories, even when they’re outside the context window.

3

⌨️ Open Interpreter — the LLM that runs code on your machine

Imagine asking it to “organize my Downloads folder by file type” and the AI simply DOES it — writing a small program, running it right then, and moving the files. That’s exactly what the Open Interpreter does: it lets the LLM write and run code directly on your computer. It’s the raw power to turn a request in everyday language into a real action on the machine.

New here? “Run code” and give the computer a list of instructions to carry out (move a file, download something, do a calculation). Usually, a programmer writes these instructions; here, AI writes them for you. That's incredibly powerful—and exactly why it's dangerous.

The safety feature that makes this acceptable is the confirmation: before running anything, Open Interpreter SHOWS you the code it’s going to run and asks, “May I?” You read it, understand what will happen, and only then approve it. This is the course’s golden rule in practice: power without a brake is an accident waiting to happen; power WITH confirmation is a useful tool.

✓ The power this unlocks

  • ✓Real tasks on the computer, not just text.
  • ✓Organizes files, does calculations, generates charts, automates repetitive tasks.
  • ✓You describe it in Portuguese; it translates that into action.

✗ The risk you take on

  • ✗Incorrect code can delete or mess up files.
  • ✗Without confirmation, a misinterpreted request can cause damage.
  • ✗NEVER approve code you don't understand.

⚠️ The lesson of “power + responsibility”

Open Interpreter is the clearest example of a principle that applies to EVERY Jarvis: the more an agent can do in the real world, the more careful the architecture needs to be. Confirmation, a sandbox (an "isolated box" where code cannot reach the rest of the system), and limits are not bureaucracy — they are what distinguish a tool from a weapon.

Key concepts

Open Interpreter

The LLM that writes and runs code on your machine.

Confirmation

It shows you the code and asks for your approval before acting.

Sandbox

An isolated box where the code can't reach the rest of the system.

Power vs. responsibility

More real-world action requires more architectural guardrails.

4

🤖 Assistants and hardware — the software is the “puppet”

So far, we’ve talked about software. But there’s another whole group out there: VOICE assistants that run in your home, and even dedicated DEVICES that promised to be “the phone of the AI era.” It’s worth getting to know both — and especially the lesson the second group left behind.

On the local voice software side, there are mature, open options: Leon e OpenVoiceOS (OVOS) (voice assistants you run at home, in the spirit of “an Alexa that’s yours”), Jan (a local AI chat app, like “offline ChatGPT”) and Home Assistant Assist (the smart home platform's built-in assistant, Home Assistant, which controls lights, locks, and sensors without sending anything to the cloud). They all share our course's thesis: privacy and control come from running locally.

📊 The story of dedicated hardware

In 2023-2024, two startups tried to sell the “AI gadget”: the Rabbit R1 (a little orange device for ~US$ 199) and Humane AI Pin (a ~US$ 699 pin that projected the screen onto your hand). Both made a lot of noise at launch — and both received a lukewarm to poor reception. Humane was even discontinued.

The issue was almost never the hardware itself. It was the EXPERIENCE: slow, wrong answers, did less than the phone you already had in your pocket.

💡 The takeaway

"The problem is rarely the hardware." What makes a good assistant is the BRAIN behind it — the model, memory, tools, personality — not the pretty little box. That's why this course teaches you to build the brain first. A good Jarvis runs on the phone you already have; you don't need a US$ 699 brooch. Later on (Track 6), when we talk about a "doll" for children, remember: the doll is just another channel — the value is in the brain.

New here? “On-device” / local = runs on the device itself, without the cloud. "Home Assistant" and a popular, free smart home platform. "Discontinued" means the company has stopped manufacturing or supporting the product — something important to remember before relying on a closed gadget.

Key concepts

Leon / OVOS / Jan

Voice and chat assistants that run locally.

Home Assistant Assist

Smart home assistant, no cloud.

Rabbit R1 / Humane Pin

Dedicated AI gadgets, with a lukewarm to poor reception.

The brain beats the box

The value is in the software, not the fancy hardware.

5

🔌 MCP — the “USB” for AI tools

Here’s the piece that makes this whole world fit together. Imagine that every tool you want to give your Jarvis (Gmail, calendar, GitHub, your notes app) spoke a different language, with a different connector. You’d have to write an adapter for every combination—a nightmare. That’s how it was until Anthropic launched, in November 2024, o MCP (Model Context Protocol). It is literally the “USB” for AI tools.

New here? Remember when every device had a different charger? USB solved that with ONE connector that works for everything. A "protocol" and it's simply an agreement about "how to communicate" that everyone accepts. MCP is the agreement that lets any tool connect to any AI agent without a special adapter.

Before USB, each peripheral came with its own cable. With USB, you plug it in and it works. MCP does the same for tools: each integration becomes a "MCP server" independent and standardized. Want to give Gmail to your Jarvis? Plug in the Gmail MCP server. Want to remove it? Unplug it. The agent doesn’t even need to be rewritten — that’s the beauty of an open standard.

NO standard · every thread is different agent A agent B Gmail GitHub Notes 6 adapters for 2x3 only WITH the MCP · a common connection point agent A agent B MCP the "USB" Gmail GitHub Notes plug in any one · without rewriting the agent

On the left, without a standard each agent needs its own thread for each tool—a web that explodes. On the right, with the MCP at the center, they all speak the same “plug-and-play language”: connect the tool server and you’re done, without rewriting Jarvis. That’s why MCP is what lets parts of the outside world connect to your home project.

🛡️ Why MCP is also SAFER

In the previous module, you saw the disaster of OpenClaw's 341 "malicious skills" — third-party files that did things behind the scenes. MCP is safer because each tool is a separate, standardized, auditable server: you know what it does, turn it on and off whenever you want, and don't need to paste random code into your agent. It's the clean way to give Jarvis hands.

Key concepts

MCP

Model Context Protocol, Anthropic’s open standard (Nov. 2024).

“USB for tools”

A single connection point that connects any tool to any agent.

MCP server

Each tool becomes a separate module that you plug in or remove.

Auditable

Safer than downloading third-party "skills" without knowing what they do.

6

🧭 THE PATH — the honest recommendation

You’ve seen a lot: papers, multimillion-dollar projects, gadgets, patterns. It’s natural to think "I need the most sophisticated option." The honest recommendation in this course is the opposite: start small, lean, and local — and grow only when a real need arises. The world out there is the map; but your first step fits into a weekend.

New here? "Lean" (lean) means starting with the MINIMUM that works, without extra weight. “Single-owner” = a system that serves ONE person (you), which is simpler and safer than one that tries to serve everyone. “Out of necessity” = only add a piece when you’ve actually missed having it — not “just in case.”

Bringing together everything you learned in Tracks 1 and 2, here’s the recipe for a safe path—each item addresses one of the dangers we saw (exposure, cost, malicious skills, complexity no one understands):

1.
Lean and local first. Start with a lean project (like Intelecto or GravityClaw) and a local model (Ollama). Fixed or zero cost, with your data at home.
2.
MCP only for tools. No copying third-party "skills." Each tool is connected through an auditable MCP server.
3.
Single-owner. Only YOU are served (your ID is whitelisted). No exposed web server, no ports open to the world.
4.
Memory in files + index. Text notes (the readable truth) with a search index on top. It survives the hype because it's just text.
5.
Grow as needed. Add voice, schedules, subagents ONLY when you miss them. Understand every building block you add.

🧱 Build by understanding every brick

The takeaway from this module: “code you understand is worth more than code you clone and never read”. The bloated OpenClaw (100 thousand lines) became a risk precisely because no one read it all. A Jarvis of a few hundred lines that you fully understand is more powerful in practice — because you trust it, fix it when it breaks, and help it grow at your own pace.

🌉 The bridge to the next tracks

Now you have the complete PANORAMA map: systems at home (module 2.2) and the world out there (this module). The next step is to open the hood and see HOW a Jarvis works inside.

  • •Track 3 (Anatomy): the 6 layers — channels, identity, tools, skills, agents, and brains — that turn a chat into an assistant that sees, remembers, and acts.
  • •Track 4 (Build): the minimum recipe and the 4-layer architecture to build your own for real, brick by brick.

Self-check (optional): Which sentence best sums up the recommended "path"?

Key concepts

Lean and local

Start small, on your own hardware, with fixed or zero cost.

Single-owner

It serves only you; no doors open to the world.

Grow as needed

Add pieces when you miss them — not just in case.

Every understood building block

Code you understand is worth more than code you clone.

🎯 Module summary

✓
AIOS — the “LLM = kernel” metaphor became an academic paper (COLM 2025), 2.1x faster at organizing agents.
✓
MemGPT / Letta — hierarchical memory (core/recall/archival) that gives the agent “infinite memory.”
✓
Open Interpreter — the LLM runs code on your machine WITH confirmation: power + responsibility.
✓
Assistants and hardware — Leon/OVOS/Jan/Home Assistant (local) and the Rabbit/Humane gadgets: the problem is rarely the hardware.
✓
MCP — the “USB for AI tools” (Anthropic, Nov/2024): a common, standardized, auditable connection.
✓
THE PATH — lean and local, MCP only, single-owner, file-based memory; grow as needed, understanding each building block.

You completed Track 2 — The Big Picture!

You now have the complete map: how to compare AI operating systems, the family at home, and the world outside. Next: open the hood in Track 3 (Anatomy) and see the 6 layers inside.