LEARNING PATH 03 / INEMA AREAS

Improve, decide, and open up

RSI, JEV, and WebMCP: improve through testing, decide through evaluation, and prepare your site for agents.

0% 0 of 0
01 / LEARNING PATH 3RSI02 / LEARNING PATH 3JEV03 / LEARNING PATH 3WebMCP04 / LEARNING PATH 3AREA WORKSHEET05 / LEARNING PATH 3FIRST STEP
Three areas, with the same deliverables for each: the worksheet and the first step.
3 areas
18 topics
60 min with practice
1 worksheet per area

Learning path map

3.1 ~20 min

🧭 RSI: AI that helps improve AI

AI improvement cycles

3.2 ~20 min

🧭 JEV: AI Decisions in Practice

AI decisions in practice

3.3 ~20 min

🧭 WebMCP: Your Site Talking to Agents

sites that communicate with agents

Detailed content

MODULE 3.1

RSI: AI that helps improve AI

Explain the three levels of improvement in AI and turn an improvement idea into a verifiable test, with a recorded version and a way to roll back.

0% 0 of 0

What it is

RSI stands for recursive self-improvement: the idea of AI helping improve AI. The area’s central question is practical: how does a system propose changes, test results, and keep what works? The page treats this as a field of study and an open collection of resources, with references, examples, and criteria for evaluating claims. The starting point is a distinction: improving a response, improving an agent, and improving the ability to create new agents are different things. An agent here is an AI system that receives a goal and carries out steps using instructions, memory, and tools.

Why learn it

Claims about AI improving itself come up often and tend to mix these three levels. If you understand the difference, you can ask what exactly improved and what evidence supports the claim.

Key concepts

recursive self-improvement; propose, test, keep; three different levels; open collection of resources; criteria for evaluating claims

What it is

The first level is reviewing the response: the model critiques and revises an output. This can improve a task without changing its parameters—the internal values learned during training—or its development process. The second level is improving the system: the agent’s instructions, memory, tools, or code change, and the versions need to be compared on held-out tasks while keeping a way to roll back. The third level is investigating recursion: the improvement helps produce future improvements. This mechanism needs evidence, because more attempts or a higher score do not prove unlimited growth.

Why learn it

Each level calls for a different kind of control. Reviewing a response is inexpensive and reversible; changing the system calls for version tracking and tests; claiming recursion calls for strong evidence. Mixing up the levels can lead you to accept guarantees no one has measured.

Key concepts

review the response; improve the system; investigate recursion; parameters; held-out tasks; ability to roll back

What it is

The page suggests starting with a small task: answer based on a policy, create questions from a text, or extract action items from notes. First define the task, the correct reference, and the limits for data, spending, and actions. Then record the initial version and separate development cases from cases held out for validation. Propose one change at a time and compare using the same criteria, including errors, cost, and review effort. Finally, validate before adopting: save results and versions, decide who has the final say, and keep a rollback path.

Why learn it

A familiar evaluation can be exploited, the page warns. Keeping tests independent, checking for invented data, and looking at results outside the sample helps you avoid mistaking a local gain for general autonomy.

Key concepts

small task; correct reference; held-out cases; one change at a time; same criteria; validate before adopting; rollback

What it is

LOOP-R is the project’s framework for organizing improvements: Run, Measure, Critique, Propose, Test, Validate, Promote, and Repeat. The page describes it as ready to use in Claude Code, with nine assistants in separate roles, a record of each version, a spending cap, and a command to roll back. The safety rule is explicit: the system cannot decide to replace the current version with a worse one. In the RSI course, LOOP-R is presented as a conceptual proposal from the project. There is also a LOOP-R course with five tracks and 21 lessons for owners and managers without a technical background.

Why learn it

A cycle without guardrails can replace a good version with a worse one without anyone noticing. A spending cap, version records, and rollback make the cycle auditable and safe to repeat.

Key concepts

Run; Measure; Critique; Propose; Test; Validate; Promote; Repeat; nine assistants; spending cap; roll back

What it is

The “Start here” section has two resources. The RSI Course v6.2 has 18 lessons in six modules, with exercises, reviews, and materials for applying LOOP-R and recording approved knowledge, in Portuguese, English, and Spanish. The RSI Guide maps the topic, including mechanisms, applications, and limits, also in three languages. The page includes an important caveat: the project brings together research and educational content; it is not a ready-to-run autonomous RSI system. The code and materials are in the rsi repository on GitHub.

Why learn it

Starting with the guide gives you the vocabulary and an understanding of the limits; the course turns that into practice with exercises. Knowing the project is educational helps you avoid expecting a tool that improves itself.

Key concepts

RSI v6.2; 18 lessons; six modules; RSI Guide; PT / EN / ES; educational content, not an autonomous system

What it is

Along with the guide and course, the resource collection includes three materials with clear caveats. RSI Copiloto is a local assistant with AI, memory, tasks, and routines that compares instructions, asks for human review, and lets you roll back changes; its public demo uses programmed examples, without AI, and real AI is available only in the local version. Dream-RSI is an independent educational guide, not affiliated with Google, about learning from experiment histories; the lab and the eight sessions are future plans, not an implementation of the paper. Alerta IA 2028 is a course with three tracks about the improvement cycle, evidence, and the limits of evaluation. It treats 2028 scenarios as hypotheses, with no guaranteed timeline. The page also lists references such as Self-Refine, Reflexion, Darwin Gödel Machine, and AlphaEvolve, reviewed on September 25, 2026; their results were not reproduced in the project.

Why learn it

Each resource says what is available now and what is still a plan. Reading those caveats is an exercise in the field itself: separating evidence from claims.

Key concepts

RSI Copiloto; demo without AI; independent Dream-RSI; Alerta IA 2028; scenarios as hypotheses; original references in English

View full version →

MODULE 3.2

JEV: AI Decisions in Practice

Explain what a structured decision with Jev is, choose between Choice, Noul, and Score, and separate what has already been observed from what still needs to be measured.

0% 0 of 0

What it is

Jev is TypeSafe’s structured decision model. A structured decision is a closed question with context and criteria. Its answer is a choice, a probability, or a level, not free-form text. The Jev Decision Lab is INEMA’s educational project for formulating questions, comparing answers, and understanding when a decision needs review. The model receives context and criteria; your system remains responsible for taking action. The lab separates this choice from text generation and tool execution.

Why learn it

Many AI tasks at work don’t call for text; they call for a decision: which queue should this ticket go to? Does this passage support the claim? Separating decisions, generation, and execution makes clear who is responsible for each part and makes it easier to measure accuracy.

Key concepts

structured decision model; TypeSafe; Jev Decision Lab; context and criteria; the system is responsible for actions

What it is

The page presents three ways to ask a question. Choice selects from explicit alternatives, such as which queue or which agent. Noul estimates the probability of an affirmative answer, such as the chance that a passage supports a claim. Score evaluates a rubric with ordered levels, such as low, medium, and high. One course module covers how to distinguish the three based on the kind of answer you need, and another covers confidence and error: using uncertainty without turning it into automatic authorization.

Why learn it

Choosing the wrong format produces answers that are hard to use: a score when you needed a queue, or a yes when there were several levels. The form of the question determines how you will measure results and which policy to apply.

Key concepts

Choice: explicit alternatives; Noul: probability of yes; Score: ordered levels; confidence is not authorization

What it is

The lab runs in your browser, requires no key, and doesn’t call any API: its answers are simulated. Each of the 20 cases includes context, questions, criteria, a simulated answer, an explanation, and a next step. Examples include support triage, a contract checklist, agent routing, and semantic diff review. You can edit the questions, import a request, export the result, or open a report; changing a case invalidates its previous simulated answer. To go further, you can clone the project and run the local server with Python 3.10 or later. The core uses only the standard library.

Why learn it

The lab lets you make mistakes without cost or risk. You learn to formulate the context, options, and policy before spending credits or touching a real system.

Key concepts

lab requires no key; simulated answers; 20 cases; context, questions, criteria; local Python server

What it is

The course "Jev in Practice" has three learning paths, 12 modules, and 36 lessons with theory, examples, exercises, and answers. The Understand path covers where Jev fits, the three question types, confidence and error, and cost. Apply covers support, documents, models and agents, and the browser. Build and Evaluate covers integration, quality measurement, operations, and the final project. Modules 1 through 8 are conceptual; modules 9 through 12 use JSON, the terminal, and Python. The estimated time is 18 hours. There are 12 labs with fictional data and answer keys. The final project asks you to decide whether there is evidence to adopt, collect more data, or not automate. The HTML v2 course is in Portuguese, with progress tracking, questions, and notes.

Why learn it

The course takes you from "understanding the idea" to "deciding whether to adopt" using a rubric. The final project accepts "do not automate" as an answer, which trains sound judgment rather than enthusiasm.

Key concepts

Understand; Apply; Build and Evaluate; 12 modules; 36 lessons; 12 labs; final project; 18 hours

What it is

A package is a kit for a specific area with context, criteria, a request, a fixture (a ready-made example for testing), and instructions. The current collection has 17 packages, covering areas from support and sales to YouTube comments, meetings, video clips, and curation. All packages share the executor and Python core: without the --live option, a package uses the simulated fixture; with it, the package makes a real query, which requires a backend key and may use credits. There is also batch execution with resume support, the jev-decidir skill for Codex and Claude Code, and jev-gw, a gateway with a daily spending cap, cache, cost logging, and a conservative failure mode that sends cases for human review. The packages are in the repository; there is no PyPI distribution, universal installer, or ready-made n8n connector yet.

Why learn it

You can move gradually: start offline, understand the output, and only then pay for real queries. The gateway shows how to protect the system that calls Jev when something fails.

Key concepts

area-specific package; fixture; --live; batch runs and resume; jev-decidir skill; jev-gw; conservative failure mode

What it is

The page separates four situations. Learning simulation: the 20 cases and replay are original and simulated; they do not measure model quality. Rules baseline: 20 correct out of 24 fictional tickets (83.33%), a result of the rules, not Jev. Real integration tested: on 09/19/2026, the ten original packages received responses through OpenRouter with the expected classifications in the fictional examples. This confirms the integration, not a benchmark; the seven new packages have only controlled tests. Evaluation still needed: quality in Portuguese, calibration, and operational use require independent data and human reference judgments. Laya, a local alternative under evaluation, passed 14 application tests and got 13 correct out of 16 synthetic examples. This does not prove it is superior.

Why learn it

This separation is what makes the area trustworthy: every number comes with what it proves and what it does not prove. Repeating these numbers out of context would overstate the results.

Key concepts

simulation; rules baseline; real integration; evaluation needed; Laya under evaluation; not a benchmark

View full version →

MODULE 3.3

WebMCP: Your Site Talking to Agents

Explain what changes when agents visit websites, what WebMCP is, and the first concrete step: assess your own site and fix the basics.

0% 0 of 0

What it is

WebMCP is a way for a page to describe what it can do to AI agents. Without WebMCP, an agent has to find fields, click, type, and interpret the screen. That breaks when the layout changes. With WebMCP, the page exposes tools with a name, description, argument schema, and response, and the agent calls the right tool by name. There are two paths: a declarative tool, where an existing form gets a toolname and a description, and an imperative tool, where JavaScript registers capabilities with JSON Schema and handles state, errors, cancellation, and fallback. This is experimental technology, currently behind an Origin Trial in the browser, which means it is available for testing.

Why learn it

It is the difference between an agent guessing and a site telling it. The site gets to decide what is allowed, what needs confirmation, and what can be canceled.

Key concepts

operating pixels vs. interacting with capabilities; declarative tool; imperative tool; JSON Schema; Origin Trial

What it is

The page lists six changes. On the visitor side: agents already use browsers to fill out forms and complete tasks; search has become asking an assistant, which creates the GEO and AEO challenge; and the standard is taking shape now. On the publisher side: a site stops being just a screen and gains a catalog of capabilities layered over the human journey; discoverability becomes a requirement, with robots.txt, sitemap.xml, llms.txt, JSON-LD, Open Graph, and a clear identity; and security stays with the site, not the model. GEO and AEO are practices for appearing in and being cited by AI assistant answers, not just search engines.

Why learn it

If you own the site, the last three changes are your to-do list. Those who understand the limitations early can learn at a lower cost.

Key concepts

agents in the browser; search has become asking; GEO and AEO; catalog of capabilities; discoverability; site security

What it is

WebMCP Readiness, at webmcp.inema.pro, opens the URL in a disposable browser, gathers observable signals, and turns the scan into a correction plan. It is a passive scanner: it checks robots.txt, sitemap.xml, and llms.txt and does not run any site tools. You get an overall score and four independent scores—WebMCP, SEO, GEO, and AEO—with evidence, alerts, blockers, up to 12 prioritized fixes, and recommended courses for each gap, in a JSON report. The scan examines only the URL you provide; it inventories the sitemap, but does not crawl its pages. The advanced scanners for each phase appear in the report as planned, not ready.

Why learn it

Starting with an assessment shows you where your site stands before you study. The page also makes the limitation clear: the scanner does not promise indexing, ranking, or citation by AI.

Key concepts

passive scanner; disposable browser; overall score + four scores; up to 12 fixes; JSON report; what requires human review

What it is

The page organizes the architecture into eight layers, in the order the training builds them. The first four describe the site: discovery (can the agent find the site?), tools (what can the site do?), contract (how do you call it and what comes back?), and state and errors (what happens when something goes wrong?). The last four protect the site: control (who confirms?), permissions (who can call?), backend and MCP (where does the truth live?), and evals and observability (is it really working?). The rule that applies across all layers: the LLM chooses the tool; the site decides whether it can run. The assessment already measures all of layer 1 and the observable signals from layers 2 and 5.

Why learn it

The layers show that exposing a tool is only the start. Without confirmation, permissions, and backend authorization, an agent can trigger something it should not.

Key concepts

discovery; tools; contract; state and errors; control; permissions; backend and MCP; evals

What it is

Before any training phase, the page lists six fixes the site owner can make, each tied to an item the scanner checks. First, run the scan and record the scores as a baseline. Second, publish the three public files: robots.txt, an updated sitemap.xml, and an llms.txt that introduces the site to AI models. Third, fix the SEO basics the agent also reads: title, description, canonical, headings in order, image alt text, Open Graph, and JSON-LD. Fourth, write for people answering questions, with direct answers, definitions, lists, and tables. Fifth, show who is behind the site, with the organization, authorship, update date, and evidence. Sixth, choose a form that does not change anything risky as your first tool, such as a search or status check.

Why learn it

These fixes do not require you to program tools, and they already raise the scores. The page sums up the method: fix things, run the scan again, and watch the score rise.

Key concepts

baseline; robots.txt, sitemap.xml, llms.txt; basic SEO; AEO: direct answers; GEO: authorship and evidence; first safe tool

What it is

WebMCP Training has an overview and four phases, all open and in Portuguese: Builder (build), Integrator (integrate), Agent Developer (orchestrate), and Expert (operate). Each phase has four modules with six topics each and uses JavaScript. If you do not code, start with the overview and diagnostic, then bring the phases to your team. Layers 1 and 2 are taught in Builder, 3 and 4 in Integrator, 5 and 6 in Agent Developer, and 7 and 8 in Expert. The page also points to AIV 2026 for AEO and GEO, the Computer Use with GPT-6 Astra course, and Super-Agentes, along with projects such as the Kit do Arquiteto de Agentes, INEMACCBOT, and os-agentes. The final invitation is direct: your own site is the best first module.

Why learn it

Knowing which phase teaches each layer lets you enter the training at the point that fits your role, instead of starting from scratch.

Key concepts

overview; Builder; Integrator; Agent Developer; Expert; four modules with six topics each; related courses

View full version →
Module complete