Durable SQLite queue, Ollama manager with RAM preflight, per-call cost with budget, and a brain that remembers in PT-BR. Runs alongside v2 until parity.

v2 works, but it's a monolith with no queue, no per-call cost tracking, and no Ollama management, and it experienced two out-of-memory crashes in August. v3 builds on what's already been completed and tested in other local repos, and organizes it in layers.
Atomic claim, heartbeat lease, backoff, idempotency, and graceful drain, ported from inemaccbot with the original 107 tests. Lanes by type of work, one claude -p at a time.
Memories with salience and decay, FTS5 + bge-m3 vector search, nightly consolidation that detects contradictions, and a curated vault that only saves with your approval.
Every LLM call goes through a single gateway: local tier β low-cost β premium, monthly budget with a hard limit, and a preflight that refuses to load a large model without 40 GB free.
Channels only translate to an event bus. The orchestrator classifies with a small local model before spending any cloud tokens.
Returns JSON with route, agent, and tier. "direto" for conversation; "agente" when it needs a file, shell, repo, web, or skill.
Budget β provider β RAM preflight β call β cost in chamadas_llm. Ollama costs nothing, but uses tokens and time.
Agent job runs claude -p --output-format json in the lane agente. Results and actual costs are sent back to the originating chat.
Every turn becomes a memory. A lasting fact generates a proposal for MEMORY.md / USER.md; you approve with a command.
Everything is local. Shared API keys are read at runtime from .env from v2; v3 only needs its own bot token.
HTTP only on 11434. Never a ollama serve in parallel.
# models used (same tags as v2) ollama pull qwen3.8:27b ollama pull llama3.2 ollama pull bge-m3
The agent runs as a CLI subprocess, not through the SDK.
node -v # v20+ claude --version # CLI in the user's PATH
Create one in BotFather. The v2 token is rejected at boot (two getUpdates = 409 and the bot goes silent).
# ~/projetos/openpcbotv3/.env TELEGRAM_BOT_TOKEN_V3=123456:AAH... PORT_V3=3142 PISO_RAM_GB=40
Five commands. The service starts even without a Telegram token, with HTTP and CLI.
Copy the example from .env and fill in only what belongs to v3.
git clone git@github.com:inematds/openpcbotv3.git && cd openpcbotv3 npm install cp .env.exemplo .env # TELEGRAM_BOT_TOKEN_V3, PORT_V3, PISO_RAM_GB, ORCAMENTO_MENSAL_USD
Checks env, Ollama and models, RAM, CLIs, active v2, unit, and memory limit. Changes nothing.
npm run doctor β same general model as v2 v2=qwen3.8:27b v3=qwen3.8:27b β RAM 63.8 GB available Β· minimum to load a large model: 40 GB β οΈ TELEGRAM_BOT_TOKEN_V3 missing β Telegram channel disabled
Snapshot with VACUUM INTO: the v2 database is never opened for writing. Idempotent.
npx tsx src/cli/importar.ts v2 memories: read 308 Β· inserted 291 Β· already existed 17
User unit with Restart=on-failure, MemoryHigh=1.5G e MemoryMax=2G. The same script can be used to restart after changes src/ or .env.
bash scripts/instalar-servico.sh # build + daemon-reload + enable + restart; no sudo journalctl --user -u openpcbotv3 -f -o cat
On Telegram, or via HTTP and CLI until there's a token.
npm run cli -- "/health" npm run cli -- "lembra que eu prefiro respostas curtas" npm run cli -- "no projeto X, conta as linhas de src/app.ts" # becomes an agent job
Queue, cost, memory, tasks, and cron without leaving Telegram. The dashboard is at http://127.0.0.1:3142/.
/status [id] # queue by lane or details for a job /usage # cost today/week/month, by tier and agent, budget /health # Ollama, RAM, RSS, queue, heartbeat, channels /memoria lista|buscar|salvar|propostas|aprovar <id> /tarefa add amanhΓ£ 9h revisar PR # reminder + summary in /daily /cron lista Β· /ollama status Β· /consolidar Β· /novo
Real exchanges from the day v3 went live, with v2 active on the same machine.
you> qual a capital do RS? e lembra que eu prefiro respostas curtas bot> Porto Alegre. bot> π Guardar no vault? "β¦prefiro respostas curtas" β /memoria aprovar 1 you> o que eu te disse que prefiro? bot> Respostas curtas.
you> no projeto openpcbotv3, conta as linhas de src/fila/worker.ts bot> π§ lead β job #7. /status 7 acompanha. bot> 419 # chamadas_llm: claude-cli Β· sonnet Β· premium Β· US$ 0,1465 Β· 7,8 s # unit memory peak: 636 MB (2 G limit)
/ollama preflight qwen3.6:35b-a3b β a large model is already resident (qwen3.8:27b) β 1-resident policy /ollama descarregar qwen3.8:27b β was not loaded by this process (may belong to v2) β rejected
GET /health
{ "ok": true, "rss_mb": 82, "canais": ["telegram","http"],
"ram": { "disponivelGb": 35.3, "swapUsadoGb": 8.5 },
"ollama": { "online": true, "carregados": ["qwen3.8:27b"] },
"orcamento": { "pct": 0, "limite_usd": 50, "travado": false } }v3 launches with its own bot alongside v2. Nothing is deleted until parity is reached: all v2 commands working in v3, 7 days with no lost jobs, and weekly costs measured.