🧠 Fundamentals
Before installing anything: why run AI locally, the vocabulary you’ll use throughout the course (LLM, agent, “AI OS”), what Ollama is, what a context window is—and the three ways to operate between private and cloud.
Read from left to right: what went to the cloud moves into the your machine, which powers memory, the agent, and the model — resulting in zero cost, offline and private.
Track map
🌍 Why local is the future
The cloud was the past
🗣️ The vocabulary
LLM, agent, AI OS
📦 What is Ollama
The key to open models
🪟 Context and parameters
Why 64k matters
⚖️ The trade-off
Privacy, performance, price
🗄️ The three modes
Vault, Connected, Cloud
Detailed content
🌍 Why local AI is the future
The industry’s direction of travel: from cloud to local. Ownership, zero cost, offline access, and the cases where this changes the game.
The industry moved everything to the cloud; now the movement is back to the user’s machine — local computing.
Understanding the “direction of travel” puts you ahead: the ability to run things locally will be as basic as using a computer.
Cloud→local, ownership, personal supercomputer (the Jensen Huang analogy).
You physically own the model and the data; nothing goes to OpenAI or Anthropic.
Stopping “renting intelligence” changes the game for cost, privacy, and control.
Ownership, data at home, no surveillance, no vendor lock-in.
After downloading the model, every use is free—there are no per-token charges or monthly fees.
Background agents can run 24/7 without giving you a bill shock.
CAPEX vs OPEX, $0/token, no rate limit.
Because the model runs on your machine, it works without a network connection—on a plane, off-grid, wherever.
Your productivity no longer depends on a connection or a service being up.
Availability, resilience, network independence.
Customer, health, and proprietary IP data, and regulated environments (SOC 2, GDPR, ISO 27001).
For many companies, local isn’t a luxury — it’s the only legal way to use AI with certain data.
Compliance, data sovereignty, the team’s “private brain.”
Local isn’t a religion: bring the best tool for the job, and switch when it’s no longer the best.
Avoids the mistake of forcing local use when the cloud delivers much more — and vice versa.
Pragmatism, “follow what works,” split by % of work.
🗣️ The vocabulary: LLM, agent, and AI OS
The terms that come up throughout the course, defined from scratch: LLM, agent, tools, memory, persona, skill, and what an "AI operating system" is.
An LLM is the type of AI model that runs behind ChatGPT—it predicts the next word based on the text.
It’s the piece you’ll download and run; knowing what it is demystifies everything else.
Model, token prediction, weights.
An agent is an LLM that uses “tools”: searching the web, running code, editing files—on its own.
Hermes is an agent; understanding tools explains why it does things, not just responds.
Tools, actions, reasoning loop.
A single place that brings together memory, skills, connections, and agents from your AI world.
And that’s what you’ll build in Track 3 — the “Hermes OS” running locally.
Orchestration, one home for everything, configurable.
Memory = what it remembers; persona = how it acts; skill = a capability you plug in.
They’re the building blocks you configure to make the agent YOUR agent.
Persistent memory, behavior, pluggable capabilities.
Integrations that give the agent access to sources — repositories, files, tools.
Connections turn the agent from a “conversationalist” into an “operator” for your work.
Integrations, external context, real-world actions.
Local = runs on your machine; cloud = runs on a company’s infrastructure.
It’s the distinction that organizes the entire course and the three modes in module 1.6.
Your own versus rented infrastructure, and where the data lives.
📦 What Ollama and open models are
The program that unlocks open models (Qwen, DeepSeek, Gemma, Mistral) on your machine — download it once and run it for free forever.
A program that downloads, manages, and runs AI models on your machine, with an app and terminal.
It’s the foundation for everything: without it, there’s no local model for the agent to use.
Local runtime, model manager, easy to use.
Models publicly released for you to download and run without asking anyone for a license.
Open source competition is what makes local computing viable and keeps making it better.
Open weights, model families, choosing by task.
The model stays on your disk; after downloading it, you only need the internet to download others.
Explains “offline” and “$0/token” in practice.
One-time download, local execution, disk cache.
It starts a local service that receives your text and returns the model’s response.
This is how the Hermes Agent will “talk” to the model in Track 2.
Local server, endpoint, model loaded in memory.
“30B” in the name means 30 billion parameters — the size of the model.
More parameters = more capable, but heavier for your machine.
Parameters, size vs. capacity, hardware cost.
With the cloud, you use their infrastructure; with Ollama, everything runs and stays on your computer.
Makes clear what you gain (privacy/cost) and what you give up (raw power).
Metered vs. local, control, trade-off.
🪟 Context window and parameters
What a context window is, what a token is, why Hermes Agent requires 64k—and how model size relates to your RAM.
How much text the model can “hold in its head” at once while responding.
It’s what limits (or enables) long tasks, such as an agent with memory.
Context, input + output limit, working memory.
The unit the model reads or generates — a token is about 3/4 of an English word.
The context window is measured in tokens; 64k tokens ≈ 25–30 thousand words.
Token, tokenization, tokens ≠ words.
Hermes Agent requires a model with at least 64,000 context tokens because of its memory and tools.
And that’s why you download Qwen 3 Coder 64k in Track 2, not just any model.
Context requirement, agent memory, tools use context.
Parameters = the size of the “brain”; context = how much it reads at once. They’re independent.
A 30B model may have a small context window; you need to look at both numbers.
Size ≠ context; check the model specs.
The model needs to fit in memory with room to spare; if it’s too large, the machine slows down.
Avoids the mistake of downloading a model that slows down your computer.
RAM/VRAM, headroom, download→test→delete.
For short conversations, a smaller, faster model is enough; 64k is for the agent.
You can have more than one model and use the right one for each task.
Fast model vs. agent model, multiple models.
⚖️ The trade-off: privacy, performance, and price
The honest truth: the best local model is about 1 year behind the frontier. What you trade off, what the benchmarks say, and how to split up your work.
With local, the data never leaves your machine — no company sees what you write.
It’s the local setup’s biggest strength and the number one reason for many use cases.
Confidentiality, data sovereignty.
The best local model today is equivalent to the best frontier model from ~12 months ago.
Set expectations: it’s very good, but it’s not the absolute best.
About a 1-year lag, in line with the pace of open source.
Beyond the hardware you already have, usage is free—no subscription.
Changes the economics of running agents all day.
$0/token, no subscription, just the cost of electricity/hardware.
Numbers comparing models (e.g., ~88,6 for Opus 4.8 vs ~74 for the Qwen you run).
Helps you read comparisons with a critical eye, without becoming beholden to benchmarks.
Benchmark, critical reading, “optimize benchmark”.
Response speed depends on your computer; large models can be slow.
You choose between speed and quality based on what you need at the time.
Latency, hardware, model size.
Imagine 100% of your work: one part calls for total privacy, another calls for maximum quality.
And it’s the reasoning that leads straight to the three modes in module 1.6.
Break tasks into parts; use the best tool for each part.
🗄️ The three modes: Vault, Connected, and Cloud
How to operate between full privacy and the best quality: Vault (private), Connected (performance), and Cloud (quality)—and when to use each one.
Vault mode: the agent only uses the local model; nothing leaves the machine.
It’s the mode for sensitive data and when you’re offline.
Vault mode, isolation, total privacy.
Connected mode, which brings more power when you need a boost.
It’s the middle ground between total privacy and maximum quality.
Performance mode, balanced.
Cloud mode, for when raw quality matters more than privacy.
Knowing when to turn on the cloud helps you avoid wasting time on local models for difficult tasks.
Cloud mode, quality > privacy, fresh web data.
Customer data, financial data, health notes, proprietary code, or simply no internet connection.
It’s the practical rule that tells you which mode to use.
Sensitivity criterion, decision rule.
You can ask Hermes to send one task to the private model and another to the cloud.
It’s the heart of Project 6 in Track 3.
Dynamic routing, “send it to private.”
Since local is free, you can leave agents working all day without usage fees.
It’s one of the biggest practical advantages of Vault mode.
Background agents, zero cost, automation.