PTENES
Library · HTTP service · CLI — no dependencies

One gateway for every decision.

A spending limit, cache, and cost logging for every Jev query. And when Jev goes down, your system doesn’t go down with it: the response becomes a request for human review.

jev-gw — one gateway for every Jev decision
What it is

Calling Jev from five different places is the problem

Each integration point reinvents error handling, nobody knows how much was spent that day, and when the API goes down, the system goes down with it. The gateway is the single entry point that solves this once.

💰 Daily spending limit

A dollar amount per day, checked before for each query. If it exceeds the limit, it becomes a human review—not a way to discover the damage afterward.

🛟 Conservative failure handling

Jev down, wrong key, timeout, unexpected response: it all returns review. No infrastructure failure is raised as an exception for the caller.

📊 Cost per call

A JSONL line with cost, latency, and decision. The state — where your customer’s data lives — isn’t recorded.

How it works

The path of a decision

Six steps, always in this order—and the order was chosen, not drawn at random.

validation→ cache→ budget→ Jev query→ policy→ log

Cache before budget

A response already in the cache costs nothing, so there’s no reason to block it because of a spending limit. Otherwise, a system that hit its limit would lose access even to what it already paid for.

Budget before the query

Checking afterward would mean checking the damage, not preventing it.

Policy after the response

Thresholds come from your domain, not the model. Email triage and clinical triage use the same Jev with different thresholds.

The document ARQUITETURA.md explains each module, the failure contract, and what the gateway no does.

Prerequisites

Python 3.10 and nothing else

No framework, no database, no queue, no pip install. The key is required only for the first actual query.

Python 3.10+

The entire gateway uses only the standard library.

python3 --version

A Jev key

TypeSafe directly or through OpenRouter. Only when making an actual query—tests and examples run without it.

export TYPESAFE_API_KEY=...
# or: OPENROUTER_API_KEY + JEV_PROVIDER=openrouter

Nothing else

Copy the folder jev_gw/ into your project if you prefer—it’s self-contained by design.

git clone https://github.com/inematds/jev-gw.git
User guide · step by step

Three ways to integrate

All three call the same function. There’s no logic that only the server has.

1

Download and run the tests

The 23 tests use a controlled evaluator—the test suite doesn’t spend a cent.

git clone https://github.com/inematds/jev-gw.git
cd jev-gw && python3 -m unittest discover -s testes
# Ran 23 tests ... OK
2

Python system: import

No extra process, no open port. It’s the most direct way to embed it.

from jev_gw import decidir

saida = decidir(pedido)
if saida['acao'] == 'suggest':
    encaminhar(saida['resposta']['answers']['fila']['choice'])
else:
    fila_de_revisao(saida['motivo'])  # includes “Jev went down”
3

Another language: start the service

Listening only on 127.0.0.1 — the gateway keeps your key.

python3 -m jev_gw servir
# jev-gw listening at http://127.0.0.1:8770 (dashboard at /painel)

curl -X POST localhost:8770/decidir \
  -H 'content-type: application/json' -d @pedido.json
4

The request

It’s the Jev API’s own format: context and one or more questions with alternatives.

{
  "model": "jev-1.13.0",
  "state": "Cliente: comprei ontem e quero devolver, nem abri a caixa.",
  "questions": {
    "fila": {
      "type": "choice",
      "instructions": "Para qual fila encaminhar?",
      "criteria": {
        "reembolso": "Pedido de devolução ou estorno",
        "suporte": "Dúvida de uso ou defeito",
        "insuficiente": "Não dá para decidir com o que está escrito"
      }
    }
  }
}
5

Adjust the limit and cache

Everything is configured through environment variables, with working defaults. 0 the gateway shuts down at the limit.

export JEV_GW_TETO_DIARIO=1.00   # dollars per day
export JEV_GW_TTL_CACHE=900     # seconds a response remains valid
export JEV_GW_TIMEOUT=5         # seconds per query
6

Track spending

A dashboard in your browser, or JSON in the terminal to add to your monitoring.

python3 -m jev_gw custo
# {"chamadas": 2, "consultas_reais": 1, "cache": 1, "gasto_usd": 1.6e-05, ...}

http://127.0.0.1:8770/painel  # daily spend, latency, latest calls
7

The failure contract

Only one error is raised as an exception: a request outside Jev’s contract, which is your bug and needs to be surfaced. Everything else is a decision.

# 200  decision made — including “review” because Jev went down
# 422  valid request, but outside Jev's contract
# 400  malformed JSON or invalid option
# exceeding the limit returns 200, not 429: for the caller, it isn’t an error,
# it is a decision to route for review
Examples

Measured on 22/09/2026, not estimated

A real Jev query through OpenRouter, followed immediately by the same query again.

Actual query

Decision reembolso, confidence 1.0.
610 ms · US$ 0.0000155

Same query, again

Served from cache.
0 ms · US$ 0.00

No key configured

Returned review logged the reason, and the client continued. No exceptions.

Raw evidence: reports/smoke-2026-09-22.jsonl. This is an integration and contract test — not a decision quality benchmark, and no savings are claimed without comparable measurements.

Roadmap

What exists and what doesn’t

Actual status as of 09/22/2026, with no promises about what hasn’t been written yet.

Ready
Library, HTTP service, and CLIThe same function powers all three. 23 tests passing, including the failure path and the end-to-end server.
Ready
Limits, cache, logging, and dashboardCache and limit measured in a real call; the dashboard shows daily spend, latency, and the latest calls.
Doesn’t
It’s not a code agent proxyIt doesn’t sit between your Codex/Claude Code and the model. That’s what jev-gateway (an independent TypeScript project) is for—it complements this project.
Doesn’t
Doesn’t execute actionsIt returns a suggestion. Your system applies it, with its own permissions. No decision here triggers an API call, email, or database write.
Next
Thresholds per questionToday, the threshold applies to the entire request; ideally, each question would have its own, because the cost of getting it wrong varies by question.
Next
Adopt in openpcbotv3The bot currently has its own gateway with a budget and observation mode. Unifying them means one place for limits, costs, and logging.