Trail map
Detailed content
🧪 Hidden instructions
How outside text gives the agent instructions, the rhyme test to see this safely, and the safeguards that hold it back.
An instruction placed inside an email, PDF, or website that the agent reads and follows. Technical name: prompt injection.
Whoever writes the text ends up controlling the agent—without breaking into anything, just by sending an email.
Hidden instructions, outside content, invisible text, change of ownership.
Email, PDF, web page, spreadsheet comment, résumé, and invoice: everything the agent reads as part of its work.
You can’t stop it from reading; you can control what it can do after it reads.
Front door, attachment, navigation, sources × tools.
A one-line message from you with a hidden line asking it to "reply in rhyme," tested in two rounds: without a rule and with a rule.
Seeing the chat itself obey is more convincing than any explanation—and there's no risk.
Rhyme test, new conversation, no connectors, two rounds.
To the model, your instructions and the email from outside arrive as one continuous stream of text, with no label showing who owns what.
Explains why no written rule eliminates risk—and why approval is still necessary.
Single queue, no labels, trained to obey, reduces but does not eliminate risk.
Read ≠ act, few tools enabled, be wary of urgency from outside and get approval before sending.
No single safeguard solves everything. Layered together, the last one catches what the others let through.
Read ≠ act, fewer tools, urgency becomes a question, approval.
A ready-to-use block for the agent's fixed instructions: outside content is information; only you can give instructions.
It's the first and cheapest barrier — as long as you know its honest limits.
Fixed instructions, warn instead of obeying, honest limits.
🗝️ Passwords and extensions
Secrets, the key file the agent can see, extensions that inherit your permissions, and the leak checklist.
Everything that proves to a system that "it's me." This includes the API key — the password to the door between programs.
Whoever has the secret gets in as you. The system doesn't verify the person, only the key.
Secret, API, API key, token.
The keys are usually stored in a configuration file, such as .env, in the folder the agent can access.
Folder access means access to everything inside it—including keys, if nobody separates them.
.env, folder access, vault, printing = data leak.
You don’t paste secrets into the chat, and the agent doesn’t show secrets on screen. Ready-to-use rule for the instructions.
The history stores everything that goes in. A secret pasted into the chat ends up living in the chat.
History, don’t paste, don’t print, [SECRET HIDDEN].
Other people’s programs that install in the browser, chat, or agent and inherit whatever permissions it has.
Installing it is like hiring a stranger to work inside your office.
Extension, its permissions, asks × promises, update.
Three questions in two minutes: who published it, what it asks for, and whether you really need it.
It prevents most problems—and many extensions do what the chat can already do without installing anything.
Who published it, list of permissions, do I really need this, test profile.
Revoke, create a new one, replace it wherever it was used, and check spending—with a checklist ready to fill out today.
Deleting the message doesn’t solve it. On the day of a leak, you don’t want to be searching for where the button is.
Revoke, create a new key, replace it, check spending.
💸 Spending, isolation, and LGPD
Spending and effort limits, a sandbox for the agent, LGPD on one page, and the four-gate checklist.
Monthly account limit, usage alert, and a virtual card just for AI services.
An agent in a loop spends money with every pass. Without a cap, you find out on the bill.
Spending limit, alert, virtual card, worst-case scenario.
Time, attempt, message, and machine limits—and what the agent does when it hits them.
A loop without a cap uses money, time, and computing resources all at once. The fix is usually one line.
Loop, maximum time, attempts, stop and notify.
Restricted folder, separate account, or sandbox (sandbox/container): an enclosed space for the agent to work in.
If it makes a mistake or gets tricked, the damage stays inside—away from your bank account and your keys.
Sandbox, container, restricted folder, copy.
Personal data, sensitive data, legal basis, purpose, necessity, and data subject rights — applied to the agent.
The law applies just the same when an AI reads the spreadsheet. You remain responsible.
Personal data, sensitive data, legal basis, purpose.
Remove columns the task doesn’t use and replace names with codes before the data reaches the agent.
The agent works the same way, and anything that leaks is worth much less.
Minimize, anonymize, key table, before and not after.
One checklist with the four gates for the agent you chose at the start of the course.
Combines Track 4 into a single document and feeds the "security" block in the Track 5 spreadsheet.
Four doors, an owner, a date, an open door = agent turned off.