1. Claim checking
Check a claim against the available evidence.
Review the wording; keep the evidence link. Don’t claim it’s universally true.
Ten practical cases for asking questions, exploring alternatives, and preparing your integration with Jev.
Public educational examples. Real inference is available through the local client with a TypeSafe credential.

Jev evaluates a context based on questions and options. This project helps you design the workflow around that response.
Support, documents, agents, and more. Each example includes alternatives, a simulated response, and a next step.
Edit the context, download the JSON, and use the CLI or local server. The key stays in the server environment.
Run local rules on a fictional dataset and review correct results, errors, and the confusion matrix.
Valid output doesn’t guarantee correct interpretation. Reviews, limits, and permissions belong to the software.
The lab does not send messages, make payments, or control external agents. Clinical, contractual, and financial cases are supervised fictional exercises.
A modern browser. The public version works without an account, key, or payment.
Python 3.10 or later and Git. The core uses only the standard library. Real queries require access to TypeSafe.
python3 --version git --version
Choose one of the twenty cases. The response is authored and simulated; changing the text requires a real request to get a new decision.
https://inematds.github.io/jev/app/
The server is restricted to your computer. Open the address shown in the terminal.
git clone https://github.com/inematds/jev.git cd jev python3 -m jev_lab serve
Contract validation does not use the API. Replace the path with the JSON file you exported.
python3 -m jev_lab validate exemplos/triagem-request.json
Configure TYPESAFE_API_KEY in the process environment or in the authorized files openpcbotv2/.env or wifi/.env. The client reads it at runtime without copying the key. The OpenRouter integration was measured across the original ten packages; this is not an independent benchmark.
python3 -m jev_lab ask exemplos/triagem-request.json
Generates metrics JSON and predictions CSV. The data is fictional, and the results measure lexical rules, not Jev.
python3 -m jev_lab batch data/tickets-sinteticos.jsonl --out reports/baseline
The tests cover contracts, policies, provider errors, and batch duplicates. No key is needed.
python3 -m unittest discover -s tests -v
Check a claim against the available evidence.
Review the wording; keep the evidence link. Don’t claim it’s universally true.
Suggest the team responsible for a request.
Suggest a reversible queue and record human corrections.
Identify points that require expert review.
Refer to a professional; no automatic legal approval.
Distinguish a proposed change from confirmation.
Highlight the message. Calendar and time zones must be handled in code.
Choose the required capability before generating.
The policy resolves the specific model name and measures final quality.
Evaluate an explicit criterion for an agent's output.
Correct once or escalate. Do not replace test execution with judgment.
Route to a specialist without expanding permissions.
Refer to the catalog; permissions remain determined by the system.
Select a candidate from the page text.
Validate the ID in the current DOM and confirm the action is allowed before clicking.
Fictional organizational study, always supervised.
Human review is required; there is no diagnosis or clinical guidance.
Organize a fictional queue according to an explicit policy.
Supervised organization; no financial orders are executed.
Explore 20 cases now: the original ten plus new practices for skills, comments, evidence, composite triage, intents, dates, logs, personal data, support intent, and diff review.
The editor supports Choice, Noul, and Score, JSON state, and multiple questions. Import a request, export the result with provenance, and open CLI reports.
python3 -m jev_lab experiment data/tickets-sinteticos.jsonl --out runs/regras
Experiment guide · Exaggerations and questions · Official models and pricing
The original ten packages are joined by inbox, YouTube comments, communities, meetings, transcript-based clips, notes, curation, and travel. The travel package evaluates each listing once and matches the answers against twelve fictional profiles using rules. The seven new examples are fictional and use combined questions.
python3 -m pacotes.executar reunioes python3 -m pacotes.qualidade reunioes python3 -m pacotes.lote reunioes data/reunioes-eventos.jsonl
The first two commands demonstrate the fixture and its simulated metrics; the third validates a batch without calling an API. Live mode requires --live and a key in the backend. Concurrency is limited and resumption is supported; platform data collection and external actions remain outside the executor.
Batches, per-question evaluation, and limitations · Choose a package
At the quoted rate of US$ 0.042 per million input tokens, 10,000 tokens cost US$ 0.00042. Ten thousand identical calls add up to US$ 4.20. Add a fallback model, human review, infrastructure, and rework.
The project brings together an analysis of Laya and references for comparing local structured decisions with Jev. The Laya adapter in Jev and the bot integration are still proposals.
In the local analysis, 14 application tests passed, and 16 previously saved responses were accepted by the Jev structural validator. The educational report shows 13 correct answers out of 16 synthetic examples; some errors have high confidence. This does not prove superiority or production quality.
The next step is to evaluate the same cases and criteria, measure total cost and latency, and maintain human review. Confidence fields should not share thresholds without validation.
Read the pilot analysis and plan · Laya INEMA · Original code · Models and model card · Laya at Eventos