Heuristic pre-submission checks for Colombian RIPS claims. Flags causal inconsistencies between a clinical note, CIE-10 diagnoses, and CUPS procedure codes, before the claim reaches MinSalud's MUV or an EPS auditor.
This service does not validate against MUV. It returns a risk assessment.
Every response carries validation_mode: "heuristic_pre_check".
The service is Django 6 + Django REST Framework, served over ASGI, with database-backed API keys, per-key rate limits, monthly quotas, a usage API, async jobs with signed webhooks, and versioned CUPS/CIE-10 catalogues.
uv sync
cp .env.example .env # then set JEV_API_KEY (or TYPESAFE_API_KEY)
uv run python manage.py migrate
DJANGO_SETTINGS_MODULE=config.settings.dev uv run python manage.py runserverIssue a key, then call the API:
uv run python manage.py create_api_key \
--organization ips-demo --name local --create-organization
# prints the token once; export it
curl -s localhost:8000/v1/rips/pre-validate \
-H "Authorization: Bearer $TOKEN" \
-H 'Content-Type: application/json' \
-d '{
"encounter_id": "ENC-2026-9042",
"patient_age": 42,
"patient_gender": "M",
"clinical_note": "Paciente masculino acude por dolor abdominal agudo en fosa iliaca derecha de 12 horas. Se realiza palpacion y ecografia abdominal.",
"procedures": [{"code": "890201"}],
"diagnoses": [{"code": "K35.8", "description": "Apendicitis aguda"}]
}' | jqRun the checks:
uv run pytest # 243 tests, offline, free
uv run pytest -m live # 11 more, real API, costs money
uv run ruff check . && uv run mypy precheck apps config
uv run python scripts/smoke.py # one end-to-end verdict, printed in full
uv run python eval/run_eval.py -v # accuracy against labelled casesThe FastAPI MVP is still present under app/ and imports the same precheck
domain package; it is retired once the Django service is deployed.
Regenerate the CUPS catalogue, and build a synthetic evaluation corpus:
uv run python scripts/build_cups_catalog.py \
tabla_1_procedimientosen_salud.pdf precheck/data/cups_2018.jsonl --report
uv run python scripts/generate_synthetic_eval.py --clean 200 --glosada 100
uv run python eval/run_eval.py --csv eval/synthetic.csv --dump results.csv| Method | Path | Notes |
|---|---|---|
| POST | /v1/rips/pre-validate |
The verdict. Sync, ~227 ms. |
| POST | /v1/rips/pre-validate/jobs |
202 + job_id; Idempotency-Key supported. |
| GET | /v1/jobs/{job_id} |
Status, and the result while its TTL lasts. |
| GET | /v1/catalog/cups/{code} |
Code lookup with provenance. |
| GET | /v1/catalog/cups?q= |
Search. |
| GET | /v1/catalog/cie10/{code} |
Code lookup. |
| GET | /v1/usage |
Metered usage for the calling organization. |
| GET | /healthz /readyz /version |
Ops. /api/schema, /api/docs for OpenAPI. |
Auth is Authorization: Bearer <api key>. Keys carry scopes
(prevalidate:run, catalog:read, usage:read), can expire, can be revoked,
and can be restricted to source CIDRs. Unauthenticated, scope-denied and
quota-exceeded requests are 401, 403 and 403 respectively; a throttled request
is 429 with Retry-After.
config/ Django project: settings/{base,dev,test,prod}, urls, asgi, celery
apps/
ops/ health, version, audit logging, error handling, request refs
accounts/ organizations, memberships, roles
api_keys/ keys, hashing, authentication, scopes, lifecycle, admin
billing/ plans and subscriptions (no payment processing)
metering/ usage events, quotas, throttling, /v1/usage, rollups
catalogdata/ versioned CUPS/CIE-10 tables, imports, provider, read API
prevalidation/ endpoint, pipeline, jobs, webhooks, admin
precheck/ the domain layer; imports no web framework
app/ the retired FastAPI MVP, on the same domain layer
precheck/ holds the judgement — catalogue resolution, the question rubric, the
upstream client, the decision table, the deterministic Spanish explanations. It
imports no web framework (enforced by tests/test_precheck_purity.py), so the
domain is testable offline and portable across edges.
The guarantee survives the rewrite, and it is structural rather than aspirational.
- Not stored, ever: the clinical note, encounter id, patient age and gender, proposed codes, explanation text. No table has a column for them.
- Stored: organization, key (hashed), usage metadata, job metadata, webhook metadata, code catalogues.
- Async exception: a job result is held in the cache for
JOB_RESULT_TTLseconds (default 900) and then discarded; a succeeded job whose result is gone returns 410 Gone. The clinical note travels to the worker in the broker message, which must be operated as ephemeral. - Logs: one metadata-only JSON line per request, from
apps/ops/audit.py. Django'sdjango.requestlogger is silenced because its tracebacks can echo submitted values, and the exception handler logs exception type and source frame only.
tests/test_zdr.py and tests/test_zdr_django.py assert all of this, including
that a secret never appears in a response and that a job row holds no clinical
field.
This is the change that mattered most. An earlier version returned a boolean
is_muv_compliant, and "the model is unsure" mapped onto false. On a
300-case corpus that produced a 30.5% false-positive rate, of which only
7.0% was the service asserting a defect; the other 23.5% was uncertainty being
reported as an accusation.
verdict |
meaning | what to do |
|---|---|---|
compliant |
no problem found, no check skipped | submit |
defect_found |
a specific defect was identified | route to a coder |
indeterminate |
the service could not tell | submit — but do not treat the claim as verified |
is_muv_compliant is retained as a deprecated alias for
verdict == "compliant". Branch on verdict. indeterminate_reasons
explains what was unclear, and findings[].detail carries Spanish prose per
finding.
POST /v1/rips/pre-validate, Authorization: Bearer <token> when CLIENT_TOKENS
is configured.
| field | required | notes |
|---|---|---|
encounter_id |
yes | Correlation only. Never logged. |
patient_age |
yes | 0–130 |
patient_gender |
yes | M | F | O |
clinical_note |
yes | Truncated at MAX_NOTE_CHARS (default 6000) and flagged. |
procedures |
yes* | Up to 25 {code, description?} line items |
diagnoses |
yes* | Up to 10 {code, description?, type?} |
proposed_cups / proposed_cie10 |
deprecated | Merged into the lists if those are empty |
* At least one of each is required; the singular fields satisfy it.
Multiple procedures is the point. A real RIPS submission reports every procedure as its own line item. An earlier single-code contract could not represent a clinical encounter plus its procedure, and every false positive that traced back to it came from exactly that shape.
CIE-10: send it. No authoritative CIE-10 table is bundled, so diagnosis
coherence cannot be checked without it. Omit it and the service returns
indeterminate with codes_unresolved naming the code.
CUPS: usually not needed. The bundled catalogue holds 7,882 codes. Send a description only for codes it lacks; it is used as a fallback, and catalogue entries always win.
{
"verdict": "defect_found",
"findings": [
{
"type": "CUPS_MISMATCH",
"severity": "high",
"detail": "La nota clinica describe imagenologia o diagnostico que no queda cubierto por los codigos CUPS enviados: 890201 (CONSULTA DE PRIMERA VEZ POR MEDICINA GENERAL). ...",
"methods": ["procedimiento_imagen_sin_codigo", "procedimiento_sin_codigo"]
}
],
"primary_finding": "CUPS_MISMATCH",
"missing_procedure_categories": ["IMAGEN"],
"unsupported_codes": [],
"indeterminate_reasons": [],
"is_muv_compliant": false,
"validation_mode": "heuristic_pre_check",
"risk_score": 9,
"risk_band": "riesgo_alto",
"issue_category": "CUPS_MISMATCH",
"explanation_hint": "...",
"suggested_cups": [],
"latency_ms": 231,
"signals": { "noul": { }, "score": { }, "ambiguous": [], "codes_unresolved": [], "checks_skipped": [] }
}| value | severity | meaning |
|---|---|---|
UNSUPPORTED_CODE |
critical | A billed code the note does not document. |
DIAGNOSIS_MISMATCH |
high | The diagnosis is not supported by the note. |
CUPS_MISMATCH |
high | The note documents a procedure the codes do not cover. |
DEMOGRAPHIC_MISMATCH |
high | The diagnosis or procedure is impossible for this patient. |
UNRESOLVED_CODE |
medium | A code could not be interpreted; a check was skipped. |
LOW_CONFIDENCE |
informational | The service could not decide. |
CUPS_MISMATCH covers several questions, so methods names which ones fired.
CODING_INCOHERENCE no longer exists: it reported a contradiction between two
mechanisms that computed the same thing, and one of them was removed.
Raw model output, unmodified, so you can retune thresholds against your own
glosa data. ambiguous, codes_unresolved, and checks_skipped are the fields
worth alerting on.
| status | meaning | what you should do |
|---|---|---|
401 |
Bad or missing token | Fix credentials. |
422 |
Malformed body | Field paths returned; values are never echoed. |
502 |
Upstream rejected the request or returned garbage | Proceed without pre-validation. |
503 |
Upstream unavailable | Proceed without pre-validation. |
Do not block a claim on this service. Correlate with support using the
X-Request-Ref response header.
A question and a finding are one-to-one. An earlier version asked both a Choice and a set of Nouls about the same subject. When they disagreed the disagreement was itself reported as a finding — 14 to 20 of 66 measured false positives, including the one genuine false positive in a held-out set. TypeSafe documents that probabilities across separate questions are not comparable, so there is no invariance to lean on. The fix was to stop asking twice.
Uncertainty is a verdict, not a soft failure — see above. 30.5% → 2.6% false positives.
Instructions are English; the clinical state is Spanish. English returned confidence 0.80 with nouls stable to three decimals; Spanish returned 0.49 with nouls drifting 0.26→0.33. TypeSafe documents English as the primary training language. The customer never sees English.
Explanations are deterministic, not generated. Jev returns typed decisions and no prose by design. Prose is assembled in code: reproducible, free, and incapable of inventing a CUPS code.
risk_score is an argmax bucket, not a measurement. TypeSafe warns against
using Score output for "the exact magnitude of a number between two levels".
A verified description is not a matching indication. The first live run
suggested CUPS 881201 — "ECOGRAFÍA DE MAMA" — for an abdominal case. Codes
are suggested only when a declared indication keyword appears in the note, and
tests/test_catalog.py pins that regression.
The unsupported-code question is scoped to procedures. Asked without that scope it fired on the specification's own golden path, naming the consultation code as unsupported because the note documented findings without narrating "a consultation occurred". It now excludes routine consultation activity, and the sentinel criterion states that a consultation counts as documented whenever the note records a history, examination, or plan.
Checks that cannot run are skipped, never guessed.
app/data/cups_2018.jsonl holds 7,882 codes extracted from MinSalud's
"TABLA 1 - PROCEDIMIENTOS EN SALUD" PDF by scripts/build_cups_catalog.py.
That PDF was produced in July 2018. CUPS is now governed by Resolución 2706
de 2025 (in force 2026-01-01, ~13,639 codes). The 2018 table is a superset in
neither direction. Every entry carries source recording this.
Replace it before production. Drop a dump of Resolución 2706 de 2025 into
the same JSONL shape ({"code": ..., "description": ...}).
missing_procedure_categories names the kind of procedure that is uncovered.
suggested_cups stays empty for most encounters: the bulk table does not encode
inclusion criteria, so the indication gate fails. Detection is unaffected — it
reasons over descriptions. To enable suggestions, populate
INDICATION_OVERRIDES in app/catalog.py.
No database, no file, no cache. One JSON line per request on stdout: timestamp,
request reference, client id, endpoint, status, latency, outcome, model, token
counts. The clinical note, encounter id, patient age, gender, proposed codes and
explanation text are never logged, at any level, including on error paths.
tests/test_zdr.py asserts this.
The end-to-end claim depends on your TypeSafe agreement. TypeSafe routes zero data retention through its legal documents as an enterprise arrangement. Confirm it is in your DPA. This service can only speak for itself.
A clinical note plus age and gender is sensitive health data under Ley 1581. Sending it to a third-party processor is a legal question for your DPA.
- Not an official MUV validator. A
complianthere is not a guarantee of acceptance. - Not a medical device. It checks billing coherence, not clinical decisions.
- Not validated on real claims. See below.
eval/run_eval.py against the 366-case synthetic corpus, with real CUPS codes
resolved from the official table:
| metric | value |
|---|---|
| compliant/flag agreement | 342/366 (93%) |
| false positives on clean claims | 7/266 (2.6%) |
| detection of constructed defects | 97/100 |
| indeterminate | 18/366 (4.9%) |
| latency | median 227 ms |
| cost | ~$0.06 per 1,000 claims |
By defect: CUPS_UNDERCODED 17/17, CUPS_WRONG_PROCEDURE 17/17,
DIAGNOSIS_MISMATCH 17/17, DEMOGRAPHIC_MISMATCH 17/17,
CODE_WITHOUT_SUPPORT 16/16, MULTI_PROCEDURE_UNDERCODED 13/16.
The five hypotheses this change was tested against, all measured rather than assumed:
| # | hypothesis | result |
|---|---|---|
| H1 | dropping the redundant Choice removes CODING_INCOHERENCE |
✓ gone |
| H2 | clean-cleared rises above 90% | ✓ 91.4% |
| H3 | detection does not regress | ✓ 97/100 |
| H4 | multi-procedure collapses the indeterminate cluster | ✓ FPs 36 → 2 |
| H5 | the unsupported-code check fires | ✓ 16/16 |
Read this with suspicion. A corpus written by the same process that wrote the rubric is a sensitivity test, not a validation. It establishes that the pipeline works on real codes and that detection does not collapse. It does not establish the false-positive rate on your claims, because real clean claims are messier than anything a generator invents.
Three iterations of this corpus each found generator bugs rather than model bugs: plan slots describing performed procedures, billing a consultation the note did not document, and an obstetric diagnosis whose stated mechanism only one note variant mentioned. Expect more.
A real evaluation needs ~200 fully-paid claims and ~100 glosada for coding reasons, with notes, codes, the auditor's reason text, and ideally whether the glosa was upheld on appeal — labelled by a certified coder.
- The CUPS table is the 2018 edition. 7,882 codes, not the current ~13,639.
- No CIE-10 table. Callers must send descriptions or diagnosis checks are skipped and reported.
suggested_cupsis usually empty until indication mapping exists.MULTI_PROCEDURE_UNDERCODEDis the weakest defect (13/16, 1 missed, 2 indeterminate) — distinguishing "the consultation is coded but the procedure is not" from "the encounter is consultation-only" is genuinely hard.- The specification's
latency_ms: 142is not achievable. Measured median is ~227 ms. - Prompt injection is mitigated, not eliminated. The note is delivered as data in a labelled field and the rubric is built from constants. TypeSafe documents that jev-1.13 "does not treat state as hostile by default". Six measured injections did not flip the verdict.
- The
Resolución 948 de 2026citation in the source specification is unverified. Confirm the governing norm before quoting it.
config/ Django project (settings, urls, asgi, wsgi, celery)
apps/ Django apps: ops, accounts, api_keys, billing, metering,
catalogdata, prevalidation
precheck/ the framework-free domain layer
catalog.py CUPS/CIE-10 catalogue, provenance, suggestion gate
config.py thresholds and credentials
schemas.py request/response contract, verdict and finding types
jev.py client protocol, live client, offline stub
questions.py the rubric — the only place a domain expert needs to look
decision.py answers to findings to a verdict
explain.py deterministic Spanish prose
data/
cups_2018.jsonl 7,882 codes extracted from the MinSalud table
app/ the retired FastAPI MVP (main, auth, logging_meta)
deploy/ Docker entrypoint, Caddyfile, release script
docs/
DEPLOYMENT.md VPS runbook
tests/ offline suite; -m live for the real API
eval/
cases.yaml hand-labelled tuning set
cases_holdout.yaml held-out set (burned; kept for regression)
run_eval.py harness
scripts/
smoke.py one end-to-end verdict, printed in full
build_cups_catalog.py regenerate the catalogue from the source PDF
generate_synthetic_eval.py build a synthetic corpus with self-validation