PTENES
Research · local AI · open source

The model observes, the code decides.

A small specialist that runs on your machine, reads each message, and answers fixed questions with yes, no, or can’t tell. Uncertain cases always go to a person.

jev-open banner: local triage specialist — the model observes, the code decides. Fixed questions; yes, no, or no evidence; rules in code; refer questions to a person; runs on CPU; open source
What it is

A Jev-style local classifier for service triage

Research and education project. It does not claim parity with Jev or production readiness. The repository examples use synthetic data, and every number is labeled [measured], [reported], or [simulated].

🧩 Three answers per question

Each question on the form (“do you want to schedule?”, “is there an urgent issue?”) becomes yes / no / uncertain, with a calibrated probability. Lack of evidence never becomes “yes.”

⚙️ Rules in code

Scheduling, deadlines, amounts, priority, and queues are handled in shared code. The model only observes, and a keyword list can escalate critical cases.

🖥️ Runs on CPU

Trains on a GPU, serves on a CPU with ONNX int8 (~570 MB), in an API without torch and in a Docker container. No message leaves the machine.

How it works

From message to the right queue

Every message still reaches a person. The specialist only prioritizes the queue and flags urgent cases.

Message→ Model: 7 questions × 3 hypotheses→ yes / no / uncertain + %→ Code: rules + keywords→ Queue + reasons

Build (one time)

Task form → reviewed examples → data → untrained baseline → train the last 2 layers → frozen final test + OOD lane.

Use (always)

API POST /classify on CPU. Empty text, text that is too long (never truncated), or an internal error goes directly to human review.

Receipts

Each stage saves a JSON with hashes, metrics, and all predictions. Nothing is overwritten, and the final test runs only once.

Prerequisites

What it needs

A GPU helps a lot for training. A CPU is enough for serving.

Python 3.12 + uv

Manages training dependencies (torch, transformers, onnx).

uv sync

Ollama (optional)

Only for generating synthetic prototype data with a local LLM, at no API cost.

ollama pull qwen3:30b

Docker (service)

The image only includes onnxruntime and tokenizers. The weights are mounted in /pacotes.

docker compose up --build
User guide · step by step

From scratch to a serving specialist

The same commands work for any niche; only the folder tasks/<nicho>/ changes. The complete step-by-step guide, including precautions for each niche, is in docs/PASSO-A-PASSO.md.

1

Clone and set up the environment

Check whether torch can see the GPU.

git clone https://github.com/inematds/jev-open && cd jev-open
uv sync
uv run python -c "import torch; print(torch.cuda.is_available())"
2

Write the niche form

Questions, three hypotheses per question, the meaning of each label, actions for each queue, a critical question, and targets. Use an existing niche as a template.

mkdir tasks/meu-nicho
cp tasks/advocacia-atendimento/{ficha.yaml,regras.py,rede_urgencia.py,__init__.py} tasks/meu-nicho/
# edit ficha.yaml, regras.py, and the keyword list
3

5 examples + smoke test

A person from the niche reviews the examples (gate 1). They count only as a test and never become training data.

uv run python tools/zeroshot_smoke.py meu-nicho   # receipt in tasks/meu-nicho/recibos/
4

Data

Real, anonymized messages are ideal. For prototyping, the synthetic data generator uses a local LLM and a second LLM as a verifier, with everything labeled as synthetic.

uv run python tools/gerar_sintetico.py meu-nicho tasks/meu-nicho/dados/sintetico.jsonl \
    --n 320 --gerador qwen3.6:35b-a3b --verificador qwen3:30b --prefixo syn
uv run python tools/pipeline.py preparar meu-nicho   # split by group + dedup + manifest
5

Baseline, training, and final test

The baseline determines whether training is worthwhile. The final test runs once, with both the baseline and trained model, on the test set and the OOD split.

uv run python tools/pipeline.py baseline meu-nicho
uv run python tools/pipeline.py treinar  meu-nicho
uv run python tools/pipeline.py testar   meu-nicho
6

Export and serve on CPU

ONNX int8, with a parity gate (±1 pp against fp32) and latency measurement. The API rejects the package if the spec has changed since training.

uv run python tools/exportar_onnx.py meu-nicho
mkdir -p pacotes && cp -r workshops/meu-nicho-v1/pacote pacotes/meu-nicho
JEV_PACOTES=pacotes uv run python -m serve.api --tarefas meu-nicho
uv run python tools/testar_servico.py http://127.0.0.1:8080 meu-nicho
Examples

Two niches in the repository

Always limited to administrative tasks: no legal opinions, no diagnosis. Keyword lists and fixed texts must be validated by professionals before any real use.

⚖️ Legal services

7 questions: schedule, service, case status, payment, document, legal question, and urgency (jail, hearing, or deadline today/tomorrow, warrant, eviction). Urgency marked “yes” or “uncertain” goes to the attorney immediately; the goal is zero missed urgent cases. Case updates only after verifying identity (confidentiality).

🩺 Clinic reception

4 questions: schedule, reschedule, document, and warning symptom. An alert marked “yes” or “uncertain” goes to a person immediately; the goal is zero missed alerts. Health data is sensitive under the LGPD, so everything runs on the clinic’s machine.

POST /classify?tarefa=advocacia-atendimento
{"texto": "Boa noite, meu filho foi preso agora há pouco e está na delegacia do centro..."}

# response format (excerpt)
{"decisao": {"prioridade": "advogado_imediato",
             "motivos": ["urgencia=sim -> advogado_imediato",
                         "rede de palavras-chave: delegacia, preso -> advogado_imediato"]},
 "observacoes": {"urgencia": {"rotulo": "sim", "confianca": ..., "probs": [...]}, ...}}
Roadmap

Where the project stands

Status as of 2026-09-25. Anything that depends on real data or a real VPS has not yet been measured.

Done
Phases 0–1: environment, forms, and smoke testtorch with CUDA on the GB10 [measured]; two niches with a spec, rules, and keywords; smoke test with no critical cases missed [measured, 5 synthetic examples].
Done
Full v1 run, with synthetic data [simulated]Law firm: test accuracy increased from 72% to 96%, with 1 missed urgent case (goal not met). Clinic: from 82% to 96%, with zero missed alerts. 570 MB ONNX int8, 0.7–1.5 s per message, and ~1.2–1.3 GB of RAM per niche with 2 threads [measured on the GB10, not on a VPS]; the API and Docker passed all 9 service tests. Details and failed criteria in RESULTADOS-v1.md.
Next
Real anonymized data + calibrated thresholdReal messages used with authorization and annotated by two people; an OOD lane from another law firm or clinic; a confidence threshold per question calibrated on the dev set; targets genuinely verified.
After
Real VPS and new tasksLatency and RAM on a VPS with 2 vCPU / 4 GB; triage of legal notices, power of attorney review, and insurance plan guides.