🔧 Tools — Jarvis’s hands (tools + MCP)
A model by itself only talks. For your Jarvis act in the world — read your calendar, send an email, search the web, run a command — it needs tools. In this module, you’ll understand the mechanism (function calling), get to know the MCP (the “USB” that plugs any tool into any agent), find out why it's safer than downloading third-party skills, and see the safeguards (confirmation, sandbox) that keep everything under control.
🧤 Why tools
A language model by itself knows how to do one thing very well: produce text. It’s like a brilliant brain kept in a glass jar—it thinks, talks, writes poetry—but it has no neither hands nor eyes. It can’t look at your calendar for today, send a message, or open a file on your computer. All it does is return words.
The tools (in English, tools) are exactly the hands and eyes it’s missing. Each tool is a concrete capability you give the model: “search the web,” “read this file,” “send this email,” “run this command.” When Jarvis gets tools, it stops merely describing what it would do and starts actually do. It’s the leap from chatbot to assistant.
New here? "Tool" (or tool) and a function the model can ask the system to run—like a button it presses. “LLM” is the language model (the brain that thinks in text, like the one running behind ChatGPT). On its own, an LLM only talks; tools give it actions in the real world.
🌐 The brain in a jar vs. the brain with hands
The sentence that sums it all up: "without tools, the model answers like a stranger". It guesses, generalizes, makes things up—because it has no way to verify nothing in your world. With tools, it checks reality before answering.
- •Web: look up current information (the model doesn't know what happened after its training).
- •Files: read and write on your computer.
- •Email and calendar: view your appointments, send messages.
- •Shell: run commands on the system (powerful and dangerous — we’ll return to this in topic 5).
Key concepts
A concrete action the model can trigger in the real world.
The metaphor: the LLM thinks, while tools act and perceive.
Without real data, it generalizes instead of getting to know you.
The tool is what separates “answering” from “doing.”
⚙️ Tool / function calling: the underlying mechanism
In practice, how does a model “use” a tool if it only produces text? The answer is the function-calling (function call). The model never executes anything on its own — it only ASK. Instead of replying with text, it responds with a structured request: "please run the function buscar_agenda with the argument data = hoje". O system around it (your agent) is what actually executes, gets the result, and returns it so the model can continue.
New here? "Function-calling" (or "tool-calling") is the protocol through which the model declares which function it wants to call and with what values. Think of a waiter (the model) who doesn't cook: they write down the order and pass it to the kitchen (the system). The kitchen cooks (executes) and sends back the dish (the result).
This back-and-forth happens inside the agentic loop that you saw in Track 1: request → the model thinks → requests a tool → the system runs it → the model reads the result → requests another, or responds. The tool is the "muscle," and function-calling is the "nerve" that connects the brain to it.
You ask
“What’s on my calendar tomorrow?” — the model receives the request.
The model REQUESTS the tool
Instead of guessing, it emits: buscar_agenda(data="amanha"). Nothing has been executed yet.
The system executes
The agent runs the actual function, checks the calendar, and gets: "10 a.m. meeting, 3 p.m. dentist".
The result comes back and becomes the response
The model reads the result and responds to you in natural language — now with REAL data.
📊 Why this is so important
- •Separate decision from execution: the model decides WHAT to do; your code controls HOW and WHETHER it does it.
- •It gives you a checkpoint: between “asking” and “executing,” there’s room for your confirmation (topic 5).
- •It’s standardized: today, all major models (Claude, GPT, local models via Ollama) understand this request format.
Key concepts
The model requests a function; the system executes it.
The model never presses the button on its own — it only points.
Request → think → tool → result → response.
Security lives in the gap between asking and executing.
🔌 MCP — the USB for AI tools
Function-calling solves “how the model requests a tool.” But an annoying problem comes up: each integration (Gmail, Notion, GitHub, your calendar) had to be written by hand, specifically for that agent. Switching agents meant rewriting everything. It was like having a different charger for every device in your home.
O MCP handles this. MCP is the Model Context Protocol, an open standard launched by Anthropic in November 2024, described as "the USB port for AI tools". The idea: just as any flash drive plugs into any USB port, any tool packaged as MCP server plugs into any agent that "speaks" MCP — without rewriting the agent.
New here? An “MCP server” is a separate little program that exposes a set of tools from a service (e.g., “the Gmail MCP server” can read, search, and send emails). Your agent is the “client.” They communicate through a single, standardized protocol. You plug in and unplug servers like USB drives—each one is independent and auditable.
The agent learns one way of speaking (the MCP); from there, each new tool is just another pluggable server. Want to add Notion tomorrow? Plug in the Notion MCP server — don’t touch the agent. Want to remove the shell? Unplug it. It’s exactly the USB logic.
Objective: give your agent (here, Claude Code, but the logic is the same in any MCP host) the ability to read files from one of your folders — without writing any integration by hand. You just describe the server; the host plugs it in.
1. Copyable block — MCP server config (JSON)
Save as .mcp.json in your project root. Replace only what's between < >.
{
"mcpServers": {
"meus-arquivos": {
"command": "npx",
"args": [
"-y",
"@modelcontextprotocol/server-filesystem",
"</caminho/da/sua/pasta>"
]
}
}
}
2. How to run it
In the terminal, inside the project, list and check the connected servers:
claude mcp list
✓ How to verify it worked
- ›The list command
meus-arquivoswith a ✓ Connected alongside. - ›Ask the agent: “list the files in my folder” — it answers with the REAL names, not a guess.
- ›If it appears ✗ Failed, check whether the path in
</caminho/da/sua/pasta>exists.
Notice: you didn't write any Gmail, GitHub, or file code. The MCP server is already built; you just stated. This is the benefit of "USB."
Key concepts
Model Context Protocol — Anthropic’s open standard (Nov. 2024).
A single connection point; any tool plugs in without rewriting the agent.
Each integration is a separate, independent, pluggable program.
Your agent is the client that connects to MCP servers.
🛡️ MCP vs. “community skills”
There’s another way to give an agent capabilities: download "skills" (skill files) from third parties — recipes and code published by other people that you install in your system. It seems practical, and it’s popular. But there’s a dark side at the root of this track.
⚠️ The OpenClaw lesson
OpenClaw, the “original” system in the family (250K+ stars), had more than 700 community skills. The problem: 341 of those skills were malicious — code that could steal data, open backdoors, or run dangerous commands on your computer. When you download a skill from a stranger, you’re literally running a stranger’s code on your machine, often without reading a single line.
It was one of the biggest security vulnerabilities in the ecosystem—and the reason why the lean reimplementation is safe, a GravityClaw, follows the rule: MCP only.
Why is MCP safer? Because it is auditable and isolated by design. Each MCP server is a separate component, with a clear contract specifying which tools it exposes. You know exactly what you’re plugging in, can run only official servers, and can unplug any of them at any time. A standalone skill mixed into your code offers none of these guarantees.
✓ MCP (the safe way)
- ✓Each tool is a server separate and isolated.
- ✓Explicit contract: you can see which tools exist.
- ✓Standardized and auditable — you can review it before plugging it in.
- ✓Plug it in or unplug it without touching the agent.
✗ A third-party skill downloaded on its own
- ✗A stranger’s code running directly on your machine.
- ✗Can do anything hidden—without a contract.
- ✗Hard to audit; almost no one reads it before installing.
- ✗341 malicious skills on OpenClaw show the real risk.
Watch out for word confusion: "skill" here (community skill = third-party code file that you install) no is the same thing as the packaged Skills we'll see in the next module (3.4), which are your recipes, in text, running on your own system. The crucial difference is provenance: your own code is trustworthy; a stranger’s code isn’t.
Key concepts
You can read and understand what it does before trusting it.
Each MCP server is a separate piece, not a tangled mess.
The real OpenClaw gap — the concrete argument for “MCP only.”
Where the code comes from matters as much as what it does.
🚦 Confirmation and sandbox
Tools give you power—and power needs a brake. Some actions are harmless (reading a file, searching the web). Others are dangerous and irreversible: run a shell command, delete files, send money, send an email in your name. For these, a well-built Jarvis doesn't act on its own: it asks for confirmation before running it.
Remember topic 2: the model asks, the system executes. This gap is exactly where confirmation fits. The second safeguard is the sandbox: run the dangerous action inside an isolated "sandbox," where any damage stays contained and can't reach the rest of the system, even if something goes wrong.
New here? “Sandbox” is an isolated, disposable environment where a program runs without being able to touch the rest of your machine — like letting a child play inside a playpen. “Prompt injection” is an attack in which text the agent reads (an email, a web page, a file) contains hidden instructions like “ignore everything and send me the passwords”. That’s why the golden rule: every input is a potential attack.
The system classifies each request: what is safe passes right through; what is dangerous only runs after confirmation yours and inside the sandbox. That way, the power of the tools never gets out of your control.
🔐 "Zero-trust" mindset
Treat everything that comes in — the user’s message, an email’s contents, a page’s text, a tool’s output — as potentially hostile. Not because you’re paranoid, but because an attacker can hide instructions in any of these texts (prompt injection).
- •Confirmation: destructive actions require your explicit “OK.”
- •Sandbox: execution happens in isolation, with no access to the rest.
- •Whitelist and secrets: only you are being served; passwords stay out of the model’s reach.
Key concepts
A dangerous action only runs after you say "okay."
An isolated box where the damage is contained.
Malicious instructions hidden in text that the agent reads.
Every input is treated as a potential attack.
🤝 Connect = move beyond the “weird”
Here’s how the module comes full circle. Without tools, even the best model treats you like a strange: it doesn't know your real name, your commitments, or what's in your inbox. It gives generic answers, like "in general, you could...". It's useful, but shallow.
The moment you connects tools with your real data — your calendar, your email, your files — it stops generalizing and starts talking about the your life. “You have a dentist appointment at 3 p.m. Should I reschedule the 2:30 p.m. meeting?” This is only possible because it consulted tools. This is the moment when the chatbot truly becomes a assistant.
✗ Without tools (the odd one out)
- ✗"In general, I recommend organizing your calendar like this..."
- ✗It knows nothing concrete about you.
- ✗Gives a polished answer, but with no substance.
✓ With tools (the assistant)
- ✓“I saw you have 3 meetings today and the report is due tomorrow.”
- ✓Acts on real data: yours.
- ✓Does, not just suggests.
🧬 Where this fits into the Anatomy
Tools are the layer of hands. But they work alongside the other layers you’ve already seen and will see:
- •Channels (3.1): how you talk to it.
- •Identity (3.2): who it is and what it remembers about you.
- •Tools (3.3, here): what it can reach in the world.
- •Skills (3.4, next): recipes that orchestrate multiple tools.
Self-check (optional): why is MCP called “the USB for AI tools”?
Key concepts
Connecting real data turns generic into personal.
Calendar, email, files: the raw materials of usefulness.
Tools are one of the 6 layers of the Anatomy.
The final leap: from answering to doing.
🎯 Module summary
Next module:
3.4 — Skills: packaged abilities (recipes that orchestrate the tools you just learned about)