PTENES
TRAIL 1

🧠 Fundamentals

Before installing anything: why run AI locally, the vocabulary you’ll use throughout the course (LLM, agent, “AI OS”), what Ollama is, what a context window is—and the three ways to operate between private and cloud.

Cloud your data leaves past Your machine the data stays at home memory agent local model $0 / token offline · private

Read from left to right: what went to the cloud moves into the your machine, which powers memory, the agent, and the model — resulting in zero cost, offline and private.

6
Modules
36
Topics
~3h
Duration
Basic
Level
Learning path progress0%
0 of 36 topics

Track map

Detailed content

1.1~30 min

🌍 Why local AI is the future

The industry’s direction of travel: from cloud to local. Ownership, zero cost, offline access, and the cases where this changes the game.

0 of 6 · 0%
What it is:

The industry moved everything to the cloud; now the movement is back to the user’s machine — local computing.

Why learn:

Understanding the “direction of travel” puts you ahead: the ability to run things locally will be as basic as using a computer.

Key concepts:

Cloud→local, ownership, personal supercomputer (the Jensen Huang analogy).

What it is:

You physically own the model and the data; nothing goes to OpenAI or Anthropic.

Why learn:

Stopping “renting intelligence” changes the game for cost, privacy, and control.

Key concepts:

Ownership, data at home, no surveillance, no vendor lock-in.

What it is:

After downloading the model, every use is free—there are no per-token charges or monthly fees.

Why learn:

Background agents can run 24/7 without giving you a bill shock.

Key concepts:

CAPEX vs OPEX, $0/token, no rate limit.

What it is:

Because the model runs on your machine, it works without a network connection—on a plane, off-grid, wherever.

Why learn:

Your productivity no longer depends on a connection or a service being up.

Key concepts:

Availability, resilience, network independence.

What it is:

Customer, health, and proprietary IP data, and regulated environments (SOC 2, GDPR, ISO 27001).

Why learn:

For many companies, local isn’t a luxury — it’s the only legal way to use AI with certain data.

Key concepts:

Compliance, data sovereignty, the team’s “private brain.”

What it is:

Local isn’t a religion: bring the best tool for the job, and switch when it’s no longer the best.

Why learn:

Avoids the mistake of forcing local use when the cloud delivers much more — and vice versa.

Key concepts:

Pragmatism, “follow what works,” split by % of work.

View Full
1.2~30 min

🗣️ The vocabulary: LLM, agent, and AI OS

The terms that come up throughout the course, defined from scratch: LLM, agent, tools, memory, persona, skill, and what an "AI operating system" is.

0 of 6 · 0%
What it is:

An LLM is the type of AI model that runs behind ChatGPT—it predicts the next word based on the text.

Why learn:

It’s the piece you’ll download and run; knowing what it is demystifies everything else.

Key concepts:

Model, token prediction, weights.

What it is:

An agent is an LLM that uses “tools”: searching the web, running code, editing files—on its own.

Why learn:

Hermes is an agent; understanding tools explains why it does things, not just responds.

Key concepts:

Tools, actions, reasoning loop.

What it is:

A single place that brings together memory, skills, connections, and agents from your AI world.

Why learn:

And that’s what you’ll build in Track 3 — the “Hermes OS” running locally.

Key concepts:

Orchestration, one home for everything, configurable.

What it is:

Memory = what it remembers; persona = how it acts; skill = a capability you plug in.

Why learn:

They’re the building blocks you configure to make the agent YOUR agent.

Key concepts:

Persistent memory, behavior, pluggable capabilities.

What it is:

Integrations that give the agent access to sources — repositories, files, tools.

Why learn:

Connections turn the agent from a “conversationalist” into an “operator” for your work.

Key concepts:

Integrations, external context, real-world actions.

What it is:

Local = runs on your machine; cloud = runs on a company’s infrastructure.

Why learn:

It’s the distinction that organizes the entire course and the three modes in module 1.6.

Key concepts:

Your own versus rented infrastructure, and where the data lives.

View Full
1.3~30 min

📦 What Ollama and open models are

The program that unlocks open models (Qwen, DeepSeek, Gemma, Mistral) on your machine — download it once and run it for free forever.

0 of 6 · 0%
What it is:

A program that downloads, manages, and runs AI models on your machine, with an app and terminal.

Why learn:

It’s the foundation for everything: without it, there’s no local model for the agent to use.

Key concepts:

Local runtime, model manager, easy to use.

What it is:

Models publicly released for you to download and run without asking anyone for a license.

Why learn:

Open source competition is what makes local computing viable and keeps making it better.

Key concepts:

Open weights, model families, choosing by task.

What it is:

The model stays on your disk; after downloading it, you only need the internet to download others.

Why learn:

Explains “offline” and “$0/token” in practice.

Key concepts:

One-time download, local execution, disk cache.

What it is:

It starts a local service that receives your text and returns the model’s response.

Why learn:

This is how the Hermes Agent will “talk” to the model in Track 2.

Key concepts:

Local server, endpoint, model loaded in memory.

What it is:

“30B” in the name means 30 billion parameters — the size of the model.

Why learn:

More parameters = more capable, but heavier for your machine.

Key concepts:

Parameters, size vs. capacity, hardware cost.

What it is:

With the cloud, you use their infrastructure; with Ollama, everything runs and stays on your computer.

Why learn:

Makes clear what you gain (privacy/cost) and what you give up (raw power).

Key concepts:

Metered vs. local, control, trade-off.

View Full
1.4~30 min

🪟 Context window and parameters

What a context window is, what a token is, why Hermes Agent requires 64k—and how model size relates to your RAM.

0 of 6 · 0%
What it is:

How much text the model can “hold in its head” at once while responding.

Why learn:

It’s what limits (or enables) long tasks, such as an agent with memory.

Key concepts:

Context, input + output limit, working memory.

What it is:

The unit the model reads or generates — a token is about 3/4 of an English word.

Why learn:

The context window is measured in tokens; 64k tokens ≈ 25–30 thousand words.

Key concepts:

Token, tokenization, tokens ≠ words.

What it is:

Hermes Agent requires a model with at least 64,000 context tokens because of its memory and tools.

Why learn:

And that’s why you download Qwen 3 Coder 64k in Track 2, not just any model.

Key concepts:

Context requirement, agent memory, tools use context.

What it is:

Parameters = the size of the “brain”; context = how much it reads at once. They’re independent.

Why learn:

A 30B model may have a small context window; you need to look at both numbers.

Key concepts:

Size ≠ context; check the model specs.

What it is:

The model needs to fit in memory with room to spare; if it’s too large, the machine slows down.

Why learn:

Avoids the mistake of downloading a model that slows down your computer.

Key concepts:

RAM/VRAM, headroom, download→test→delete.

What it is:

For short conversations, a smaller, faster model is enough; 64k is for the agent.

Why learn:

You can have more than one model and use the right one for each task.

Key concepts:

Fast model vs. agent model, multiple models.

View Full
1.5~30 min

⚖️ The trade-off: privacy, performance, and price

The honest truth: the best local model is about 1 year behind the frontier. What you trade off, what the benchmarks say, and how to split up your work.

0 of 6 · 0%
What it is:

With local, the data never leaves your machine — no company sees what you write.

Why learn:

It’s the local setup’s biggest strength and the number one reason for many use cases.

Key concepts:

Confidentiality, data sovereignty.

What it is:

The best local model today is equivalent to the best frontier model from ~12 months ago.

Why learn:

Set expectations: it’s very good, but it’s not the absolute best.

Key concepts:

About a 1-year lag, in line with the pace of open source.

What it is:

Beyond the hardware you already have, usage is free—no subscription.

Why learn:

Changes the economics of running agents all day.

Key concepts:

$0/token, no subscription, just the cost of electricity/hardware.

What it is:

Numbers comparing models (e.g., ~88,6 for Opus 4.8 vs ~74 for the Qwen you run).

Why learn:

Helps you read comparisons with a critical eye, without becoming beholden to benchmarks.

Key concepts:

Benchmark, critical reading, “optimize benchmark”.

What it is:

Response speed depends on your computer; large models can be slow.

Why learn:

You choose between speed and quality based on what you need at the time.

Key concepts:

Latency, hardware, model size.

What it is:

Imagine 100% of your work: one part calls for total privacy, another calls for maximum quality.

Why learn:

And it’s the reasoning that leads straight to the three modes in module 1.6.

Key concepts:

Break tasks into parts; use the best tool for each part.

View Full
1.6~30 min

🗄️ The three modes: Vault, Connected, and Cloud

How to operate between full privacy and the best quality: Vault (private), Connected (performance), and Cloud (quality)—and when to use each one.

0 of 6 · 0%
What it is:

Vault mode: the agent only uses the local model; nothing leaves the machine.

Why learn:

It’s the mode for sensitive data and when you’re offline.

Key concepts:

Vault mode, isolation, total privacy.

What it is:

Connected mode, which brings more power when you need a boost.

Why learn:

It’s the middle ground between total privacy and maximum quality.

Key concepts:

Performance mode, balanced.

What it is:

Cloud mode, for when raw quality matters more than privacy.

Why learn:

Knowing when to turn on the cloud helps you avoid wasting time on local models for difficult tasks.

Key concepts:

Cloud mode, quality > privacy, fresh web data.

What it is:

Customer data, financial data, health notes, proprietary code, or simply no internet connection.

Why learn:

It’s the practical rule that tells you which mode to use.

Key concepts:

Sensitivity criterion, decision rule.

What it is:

You can ask Hermes to send one task to the private model and another to the cloud.

Why learn:

It’s the heart of Project 6 in Track 3.

Key concepts:

Dynamic routing, “send it to private.”

What it is:

Since local is free, you can leave agents working all day without usage fees.

Why learn:

It’s one of the biggest practical advantages of Vault mode.

Key concepts:

Background agents, zero cost, automation.

View Full