PTENES
MODULE 2.5

🧠 Memory and Context

A Jarvis without memory starts from scratch with every message — and every conversation becomes a tiring reintroduction. Here you’ll learn what the system remembers, for how long, what informs each decision, and how to retrieve the right information at the right time — without turning it into a snooping archive. Good memory isn’t about storing everything: it’s about storing what supports continuity.

Input what the user says short-term memory long-term memory Retrieve only the right context Decision on-point response Memory supports recovery; recovery supports decision-making — storing isn't the end; remembering at the right time is.
6
Topics
~45
Minutes
Intermediate
Level
Concept
Type
Module Progress0 of 6 · 0%
1

📌 What Jarvis Remembers

By default, an AI model doesn't remember anything: each message arrives as if it were the first. Memory is what you add around the model so it doesn't start from scratch each time. It's the difference between a support agent that recognizes you and already knows your usual order, and one that asks your name every time. Memory isn't model magic—it's an architecture decision made by the intent architect.

🧩 Memory is everything that survives between messages

Without memory, Jarvis lives in an eternal present: each turn is isolated. With memory, it accumulates—what was said, who the person is, what has already been resolved. The architect decides what’s worth carrying forward and for how long. Remembering takes space and attention, so it remembers deliberately, not by accident.

✓ With memory, Jarvis

  • ✓Recognizes who’s speaking.
  • ✓Picks up where you left off.
  • ✓Doesn’t repeat questions that have already been answered.
  • ✓Sounds coherent over time.

✗ Without memory, Jarvis

  • ✗Asks your name with every message.
  • ✗Forget what you just said.
  • ✗Contradicts what it said before.
  • ✗Treats every turn as a brand-new stranger.
Pattern
the model doesn't remember
Memory
layer around it
Decision
of architecture
Effect
continuity
2

🗂️ Types of Memory

Not all memory is the same. There are three time horizons the architect needs to distinguish: a session memory (lasts only for the current conversation), the short-term memory (lasts for a few interactions or for the day) and the long-term memory (carries across conversations — what defines who that person is). Mixing the three is one of the biggest sources of confusion in AI projects.

⏱️ The three memory time frames

  • Session: what's being said now, in this conversation. It ends when the conversation ends — it's the notepad on the table.
  • Short term: what's worth remembering for a few interactions or for the day ("you already gave me this request today"). It expires on its own.
  • Long term: what defines the person and carries across all conversations — name, preferences, relevant history. It's the persistent profile.

How memory evolves during a conversation

1. Customer opens the conversation

Long-term memory is consulted: Jarvis already knows the name and history. The session starts empty.

2. During the Conversation

Each message goes into session memory. Jarvis connects "the product I mentioned" to the item referenced three turns ago.

3. Something worth saving comes up

"I prefer to be called by my nickname" promotes a session fact to long-term memory.

4. The conversation ends

Session memory is discarded; only what was promoted survives for the next conversation.

Session
just this conversation
Short term
hours or the day
Long term
runs through everything
Promote
session → long
3

🍱 Context: what informs the decision

Memory is what Jarvis has stored; context is what it puts on the table to make a decision now. They’re different things. With each response, the system puts together a “plate” with pieces of memory, the current message, the identity (the soul), and the rules. That plate is the context—and the quality of the response depends much more on what goes into it than on the model’s size.

💡 Practical tip

When Jarvis “answers incorrectly,” the first question isn’t “which prompt should I improve?” but "what was in the context when this response was given?". Many failures come from missing context (it didn’t receive the information) or excess context (noise that distracted it)—not a lack of intelligence.

What typically makes up the context for a response

•The soul: who it is and its limits.
•The current message: what's being requested.
•Memory passages: only the relevant ones.
•Business data: price, inventory, status.
•The rules: what's allowed and what's not.
•The task in progress: where we are in the process.
Memory
what it stores
Context
what goes in now
Quality
comes from context
Diagnosis
"what was on the table?"
4

🔎 Retrieve the Right Information

Storing a lot doesn't help if Jarvis can't find the right piece at the right time. Retrieval is the art of bringing back only the passages relevant to the current question from a large memory. You don’t pour the entire memory into the context—that’s expensive, slow, and gets in the way. You retrieve by relevance: what relates to what’s being asked now.

🎯 The golden rule of retrieval

Bring enough to answer well, and nothing more. Too much memory in the context doesn’t make Jarvis smarter—it makes it more confused and more expensive. The secret isn’t remembering everything; it’s remembering what’s useful for this response.

✓ Healthy recovery

  • ✓Brings back passages related to the current question.
  • ✓Prioritize what's most recent and relevant.
  • ✓Summarize long content before using it.
  • ✓Leave out what isn’t relevant.

✗ Poor recovery

  • ✗Puts the entire memory into the context.
  • ✗Mixes unrelated topics.
  • ✗Repeats old, outdated data.
  • ✗Costs a lot and responds slowly.

Retrieval in one line of reasoning

question: "what was the delivery time for my order?"
search_for: requests + this client + deadline
brings: just order #4471 and its deadline
ignores: old conversations about other topics
final_context: lean and precise

Illustrative relevance-based retrieval diagram — not real code.

Relevance
connected to the question
Recency
the newest one carries more weight
Summarize
compress the long version
Sufficient
not dump everything
5

🚫 What NOT to store

A beginner’s temptation is to save everything “just in case.” The architect does the opposite: decides what NOT to store. Two reasons matter here—privacy (sensitive data that shouldn’t be retained) and noise (junk that only gets in the way of retrieval). Storing everything is both a legal risk and poison for quality. Clean memory is useful memory.

✗
Sensitive data without a need — privacy risk.

Passwords, complete documents, card details: if you don’t need to store them, don’t.

✗
Chatter and noise — the poison of retrieval.

"Good morning," "thank you," venting without facts: they take up space and make it harder to find what matters.

✗
Information that becomes outdated quickly — becomes a lie.

Today's price, current inventory: storing them as fixed facts makes Jarvis state outdated information.

💡 Practical tip

Before storing anything in long-term memory, ask three questions: is this sensitive? (if yes, avoid or protect it), will this be useful again? (if not, discard it) and will this still be true tomorrow? (if not, look it up at the source instead of saving it).

Privacy
don't store sensitive data
Noise
outside of memory
Validity
what gets outdated
Clean memory
useful memory
6

🔗 Memory in Service of Continuity

All of this exists for one reason: continuity. Memory isn’t a trophy for how much the system accumulates—it’s what makes Jarvis feel like one coherent presence instead of a thousand disconnected agents. The ultimate goal isn’t “to remember a lot”; it’s to make the person feel they’re talking to someone who knows it and picks up where they left off.

🧭 The compass of memory

When faced with “Should I save this?”, the right question is: will this help Jarvis continue the relationship better tomorrow? If it helps, it's worth remembering. If it doesn't, it's baggage. Continuity is the criterion that separates memory from accumulation—and it's what turns isolated responses into a single, reliable experience.

The contrast in one line each

sem_memoria: "Hello! How can I help?" (again)
com_memoria: "Hi, Ana! About yesterday's order..."
blind_accumulation: remembers everything, finds nothing, leaks data
memoria_boa: remembers what matters, at the right time

Illustration of how memory affects the experience — not an actual configuration.

Objective
continuity
Criterion
"can you help tomorrow?"
Presence
unique and consistent
Outcome
relationship, not turns

Self-check (optional): what criterion separates good memory from mere accumulation?

🧠 Module summary

✓
What Jarvis remembers — memory is the layer around the model, an architectural decision.
✓
Types of memory — session, short-term, and long-term; each with its own lifespan.
✓
Context — what goes on the table to make a decision now; quality starts with it.
✓
Retrieve the right thing — retrieve enough to be relevant, and nothing more.
✓
What NOT to store — privacy and noise; clean memory is useful memory.
✓
In service of continuity — the goal isn’t to remember a lot; it’s to continue the relationship.

Next module:

2.6 — Safety Boundaries: how Jarvis operates within clear, safe limits.