PTENES
Personal assistant Β· Telegram Β· local

A bot that works without exceeding the machine's memory

Durable SQLite queue, Ollama manager with RAM preflight, per-call cost with budget, and a brain that remembers in PT-BR. Runs alongside v2 until parity.

Illustration of openpcbot v3: personal assistant running locally
What it is

Successor to openpcbot v2, designed from real pain points

v2 works, but it's a monolith with no queue, no per-call cost tracking, and no Ollama management, and it experienced two out-of-memory crashes in August. v3 builds on what's already been completed and tested in other local repos, and organizes it in layers.

🧱 A real queue

Atomic claim, heartbeat lease, backoff, idempotency, and graceful drain, ported from inemaccbot with the original 107 tests. Lanes by type of work, one claude -p at a time.

🧠 Brain in PT-BR

Memories with salience and decay, FTS5 + bge-m3 vector search, nightly consolidation that detects contradictions, and a curated vault that only saves with your approval.

πŸ’Έ Cost and RAM under control

Every LLM call goes through a single gateway: local tier β†’ low-cost β†’ premium, monthly budget with a hard limit, and a preflight that refuses to load a large model without 40 GB free.

How it works

How a message flows

Channels only translate to an event bus. The orchestrator classifies with a small local model before spending any cloud tokens.

Telegram / CLI / HTTP→ Bus→ Router (llama3.2)→ Memory (3 layers)→ Direct response (qwen3.8) or agent job (claude -p)→ Secrets guard→ Chat
1

Router

Returns JSON with route, agent, and tier. "direto" for conversation; "agente" when it needs a file, shell, repo, web, or skill.

2

LLM gateway

Budget β†’ provider β†’ RAM preflight β†’ call β†’ cost in chamadas_llm. Ollama costs nothing, but uses tokens and time.

3

Queue

Agent job runs claude -p --output-format json in the lane agente. Results and actual costs are sent back to the originating chat.

4

Learning

Every turn becomes a memory. A lasting fact generates a proposal for MEMORY.md / USER.md; you approve with a command.

Prerequisites

What needs to be live

Everything is local. Shared API keys are read at runtime from .env from v2; v3 only needs its own bot token.

Ollama as a service

HTTP only on 11434. Never a ollama serve in parallel.

# models used (same tags as v2)
ollama pull qwen3.8:27b
ollama pull llama3.2
ollama pull bge-m3

Node 20+ and Claude Code

The agent runs as a CLI subprocess, not through the SDK.

node -v          # v20+
claude --version # CLI in the user's PATH

Your own Telegram bot

Create one in BotFather. The v2 token is rejected at boot (two getUpdates = 409 and the bot goes silent).

# ~/projetos/openpcbotv3/.env
TELEGRAM_BOT_TOKEN_V3=123456:AAH...
PORT_V3=3142
PISO_RAM_GB=40
User guide Β· step by step

From clone to a responsive bot

Five commands. The service starts even without a Telegram token, with HTTP and CLI.

1

Install and configure

Copy the example from .env and fill in only what belongs to v3.

git clone git@github.com:inematds/openpcbotv3.git && cd openpcbotv3
npm install
cp .env.exemplo .env   # TELEGRAM_BOT_TOKEN_V3, PORT_V3, PISO_RAM_GB, ORCAMENTO_MENSAL_USD
2

Run the doctor

Checks env, Ollama and models, RAM, CLIs, active v2, unit, and memory limit. Changes nothing.

npm run doctor
βœ… same general model as v2   v2=qwen3.8:27b v3=qwen3.8:27b
βœ… RAM                          63.8 GB available Β· minimum to load a large model: 40 GB
⚠️  TELEGRAM_BOT_TOKEN_V3        missing β€” Telegram channel disabled
3

Import memories from v2

Snapshot with VACUUM INTO: the v2 database is never opened for writing. Idempotent.

npx tsx src/cli/importar.ts
v2 memories: read 308 Β· inserted 291 Β· already existed 17
4

Install the service

User unit with Restart=on-failure, MemoryHigh=1.5G e MemoryMax=2G. The same script can be used to restart after changes src/ or .env.

bash scripts/instalar-servico.sh
# build + daemon-reload + enable + restart; no sudo
journalctl --user -u openpcbotv3 -f -o cat
5

Chat

On Telegram, or via HTTP and CLI until there's a token.

npm run cli -- "/health"
npm run cli -- "lembra que eu prefiro respostas curtas"
npm run cli -- "no projeto X, conta as linhas de src/app.ts"   # becomes an agent job
6

Operate through chat

Queue, cost, memory, tasks, and cron without leaving Telegram. The dashboard is at http://127.0.0.1:3142/.

/status [id]        # queue by lane or details for a job
/usage              # cost today/week/month, by tier and agent, budget
/health             # Ollama, RAM, RSS, queue, heartbeat, channels
/memoria lista|buscar|salvar|propostas|aprovar <id>
/tarefa add amanhΓ£ 9h revisar PR   # reminder + summary in /daily
/cron lista Β· /ollama status Β· /consolidar Β· /novo
Examples

What happened during the first live tests

Real exchanges from the day v3 went live, with v2 active on the same machine.

Memory across sessions

you> qual a capital do RS? e lembra que eu prefiro respostas curtas
bot>  Porto Alegre.
bot>  πŸ“ Guardar no vault? "…prefiro respostas curtas" β†’ /memoria aprovar 1
you> o que eu te disse que prefiro?
bot>  Respostas curtas.

Agent job with actual cost

you> no projeto openpcbotv3, conta as linhas de src/fila/worker.ts
bot>  🧠 lead β€” job #7. /status 7 acompanha.
bot>  419
# chamadas_llm: claude-cli Β· sonnet Β· premium Β· US$ 0,1465 Β· 7,8 s
# unit memory peak: 636 MB (2 G limit)

RAM preflight refusing

/ollama preflight qwen3.6:35b-a3b
β›” a large model is already resident (qwen3.8:27b) β€” 1-resident policy

/ollama descarregar qwen3.8:27b
β›” was not loaded by this process (may belong to v2) β€” rejected

Health consumed by the hub

GET /health
{ "ok": true, "rss_mb": 82, "canais": ["telegram","http"],
  "ram": { "disponivelGb": 35.3, "swapUsadoGb": 8.5 },
  "ollama": { "online": true, "carregados": ["qwen3.8:27b"] },
  "orcamento": { "pct": 0, "limite_usd": 50, "travado": false } }
Roadmap

Throttle, don't cut off abruptly

v3 launches with its own bot alongside v2. Nothing is deleted until parity is reached: all v2 commands working in v3, 7 days with no lost jobs, and weekly costs measured.

0–2 βœ“
Foundation, queue, Ollama, cost, telemetryBus, YAML config, unit with a memory limit; ported queue; Ollama manager with preflight; LLM gateway with budget; /status, /usage, /health, alerts.
3–5 βœ“
Channels, orchestrator, brain, tasksTelegram/CLI/HTTP; local router; CLI agent with session and actual cost; PT-BR memory imported from v2, embeddings, nightly consolidation, curated vault; cron as jobs, heartbeat, /tarefa, /daily.
6–7 βœ“
Slack, WhatsApp, dashboard, doctor, backupSlack via Web API (disabled until cutover); WhatsApp with a guard that refuses while v2 owns the session; single-page dashboard; doctor; encrypted nightly backup.
8
CutSwitch the production token, keep v2 read-only for 30 days, then archive it. Only with explicit authorization. Afterward: real WhatsApp in a separate daemon, enable Slack, set job table retention.