Skip to content

Repository files navigation

decision-log

An agent that reviews EU public procurement awards against their published criteria, pauses for a human to sign off, and leaves a record you can reopen months later.

The agent is not the point. The record is.

Contents. The problem · A worked example · The one idea the rest depends on · How each part works · Stack · Quick start · Commands · Testing · What this is not · What changes at scale · Status

The problem

A city publishes a call for tenders. The notice says the winner is chosen on price, weighted 80, and quality, weighted 20.

Three weeks later the city publishes a correction. Nothing unusual, it happens constantly. Six months after that the contract is awarded, and a company that lost asks which criteria its bid was actually judged against.

If software helped review that file, you need an answer. And the documents have changed in the meantime, which is where it gets hard: did the software miss the correction, or had the correction not been published yet when it looked? Without a record of what it was looking at, you cannot tell the two apart.

That is the question this repository answers, for one decision at a time.

A worked example

Everything below uses one real procurement, so you can follow the same steps and get the same output. Procedure 3978f4f3-4130-472a-831a-3eb7889b4470 is a Finnish catering contract. TED holds three publications for it.

$ decision-log ingest --procedure 3978f4f3-4130-472a-831a-3eb7889b4470
3978f4f3-4130-472a-831a-3eb7889b4470: 3 new revisions

Those three documents, oldest first:

Published Kind Publication Award criteria
2023-10-13 contract notice 622805-2023 price 80, quality 20
2023-11-02 corrigendum 666684-2023 price 80, quality 20
2024-01-03 award notice 3869-2024 price 80 Kokonaishinta, quality 20 Laatu

The middle row is the interesting one. It is a correction that restates the criteria, so it replaces the original notice. Any answer about "the published criteria" has to come from 666684-2023, not from 622805-2023.

Now review it, as the paperwork stood on 1 June 2024:

$ decision-log run --procedure 3978f4f3-4130-472a-831a-3eb7889b4470 \
    --as-of 2024-06-01T00:00:00+00:00
run 108d72bd-89b7-4888-9ddb-d2b6163bfe31 is awaiting approval

It stopped. A person has to decide, and until they do, nothing is recorded as a decision. Marie approves it:

$ decision-log approve --run 108d72bd-89b7-4888-9ddb-d2b6163bfe31 \
    --as marie@example.org --decision approved \
    --rationale "The corrigendum restates price 80 / quality 20 and the award notice matches."
run 108d72bd-89b7-4888-9ddb-d2b6163bfe31 approved by marie@example.org

Months later, this is the payoff:

$ decision-log replay 108d72bd-89b7-4888-9ddb-d2b6163bfe31

run          108d72bd-89b7-4888-9ddb-d2b6163bfe31
as of        2024-06-01 00:00:00+00:00
agent        2296da5db77152dbdaea94e2943aebafc9f72d0a / rules / rules
chain        3 entries, intact
citations    6 verified, no errors
scorecard    1
spans        5
revisions as they stood:
  2023-10-13  contract_notice        622805-2023    v1
  2023-11-02  corrigendum            666684-2023    v1
  2024-01-03  award_notice           3869-2024      v1
verdict      consistent
summary      2 criteria compared, 0 mismatched, across 3 revisions.
approved by  Marie Dubois <marie@example.org> on 2026-09-07: The corrigendum
             restates price 80 / quality 20 and the award notice matches.

Every line is carrying something:

  • revisions as they stood lists the documents that existed on 1 June 2024, not the documents that exist today.
  • chain 3 entries, intact is a check that nobody edited the log afterwards.
  • citations 6 verified means the six quoted pieces of evidence still match the documents, character for character.
  • scorecard 1 is the test run that was passing when the decision was made.
  • approved by is a named person and their reasoning, with a timestamp.

decision-log replay <run-id> --bundle out.html writes the same thing as a document laid out under the nine headings of Annex IV of the EU AI Act.

The one idea the rest depends on

Replay reconstructs. It never re-runs the model.

The tempting design is: to find out what the AI decided, run the AI again on the same inputs. That is wrong twice over. Language models are not deterministic, so you would get a different answer and no way to tell a bug from ordinary variation. And even with a perfectly repeatable model, running it today tells you what today's system does, while the question was about that system on that day.

So every step writes its output into rows that cannot be edited, and replay is careful reading. No model is ever invoked during a replay.

How each part works

1. Ingestion

What. Downloads procurement notices from TED, the EU tenders database, and stores every publication as a row that is never updated.

How. TED groups a procurement's paperwork under a procedure-identifier. All three documents in the example above share one. The idempotence key is a hash of the response content, so anything that changed becomes a new revision and an unchanged re-fetch does nothing. Running ingest twice is safe.

Why content and not a version number. TED both republishes a notice under a new number and overwrites one in place, and it only indexes the current version. A hash makes no assumption that the source is honest or consistent about announcing its own changes. For an audit record, "store what we saw, when we saw it" is a stronger guarantee than "store what the source says changed."

TED's links field, a per-language block of PDF URLs, is excluded from that hash. It is navigation rather than content, and hashing it would turn any change to TED's link structure into a phantom revision of every notice.

2. as_of

What. Every run is pinned to a date and can only see documents published on or before it.

How. One SQL clause:

where notice_id = %s
  and published_at <= %s

Why it matters. This is the whole of "the documents as they stood that day". No snapshots, no time-travel database, one predicate over rows that never change. The example above run with --as-of 2023-11-01 would see only the original notice and report insufficient_evidence, because the award notice did not exist yet. Same code, same database, different point in time.

3. Cited evidence

What. Every claim the system makes points at a specific piece of a specific document.

How. A citation stores the revision, the segment, and the SHA-256 the text had at the time. The document is split into the smallest units worth quoting, each hashed:

Award criterion 1: type price; name unstated; weight 80 (poi-exa).

Why a pointer and not a copy. Copied text proves nothing, because anyone could have typed it. A pointer plus a hash is checkable: replay looks the segment up, re-hashes it, and either it matches or it fails loudly. A finding with no citation cannot even be constructed, because the schema requires at least one. An uncited claim in an audit record is worse than no record, since it looks like evidence and is not.

4. The agent, and the pause

What. Four steps: gather the evidence, assess it, wait for a human, record the decision.

How. The waiting step calls LangGraph's interrupt(). The graph stops mid-function and its state is written to Postgres. The process can be killed, the container rebuilt, the machine rebooted. Days later someone approves, and execution resumes from inside that function as if nothing happened.

Why a graph framework at all. Only for that. Writing it yourself means designing a state machine, a serialisation format and a resume protocol. Nothing else here needs LangGraph, and the README should say so rather than implying a broader dependency.

A test proves it: the connection is closed, the pool is closed, every cache is cleared, and only then does the run resume and complete.

5. The ledger

What. An append-only log where each step of every run is recorded.

How. Each row stores a hash of its own contents and the hash of the row before it:

seq step output hash prev hash entry hash
1 gather a3f2... 0000... 7d1c...
2 assess 9b0e... 7d1c... e4a8...
3 record 1f77... e4a8... 55b2...

Why chain them. It makes the rows interdependent. Change row 1 and its hash changes, so row 2's prev hash no longer matches, and so on down the file. You cannot quietly edit one detail.

This is a hash chain: rows that commit to their predecessor. There is no network, no consensus and no distributed ledger involved, and it is worth being precise about that so nobody arrives expecting one.

One implementation note. seq is assigned by hand under an advisory lock rather than by a database sequence, because the hash covers the row's position, so the position has to be known before the insert and a sequence assigns it during.

6. Two database roles

What. The application connects with an account that cannot rewrite the record.

How. Eight lines of SQL:

grant select, insert on all tables in schema public to dlog_app;
grant update (status) on run to dlog_app;

revoke update, delete on notice_revision, chunk, ledger_entry,
                         approval, eval_scorecard from dlog_app;

A separate account owns the schema and is used only by make migrate. The one exception is run.status, because a run legitimately moves from running to awaiting_approval to approved, and it is a column-level grant rather than a table-level one.

Why bother, when the code simply does not write those statements. Because "our code does not do that" is a promise and "the credential cannot do that" is a fact. A bug, a bad merge or an injection somewhere gets a permission error instead of a silent rewrite. A test connects as the application role and confirms UPDATE on the ledger is refused.

7. Chain anchoring

What. decision-log anchor writes the chain's current tip to a file outside the database.

How. It appends {at, seq, entry_hash} to data/anchors/heads.jsonl, and refuses to anchor a chain that does not already verify. verify-chain then checks every recorded tip against the ledger.

Why. Without it, someone with database superuser access could rewrite a row and regenerate every hash after it, and the chain would verify cleanly. With it, they would have to edit the database and the file. That is a higher bar, not an impossible one, and the honest limit of what a single host can offer.

An anchor file belongs to the ledger it anchors. Resetting a development database leaves the anchors pointing at rows that no longer exist, and verify-chain will say so until you remove them; make reset discards both together. In production that pairing is the point: you cannot drop the database and quietly carry on.

8. The evaluation gate

What. A run cannot start unless the system's tests were passing.

How. decision-log eval scores a labelled set and writes one eval_scorecard row. run.scorecard_id is a not null foreign key pointing at it.

Why a foreign key. Because it is then not a policy document, a checklist item or a CI step someone can skip. The database refuses the insert. Five properties are scored, and two of them have to be perfect:

Metric Threshold What it means
Verdict accuracy 0.80 Did it reach the right conclusion
Provenance accuracy 0.95 Did it read the document it was supposed to read
Amendment recall 0.90 On cases where a correction changed the criteria, both of the above
Citation validity 1.00 Every quote resolves and still hashes to what was recorded
as_of isolation 1.00 No run cited a document published after its own cut-off

A model getting a hard judgement call wrong is a disclosed limitation. A model inventing a citation, or reading a document that did not exist yet, is a fabricated record. That is why the last two are invariants and not accuracy scores.

The gate also refuses a labelled set with fewer than eight amended cases, because a recall figure over one or two cases is a number rather than a measurement.

9. Where the labelled set comes from

What. 26 cases, 16 taken verbatim from TED and 10 built from real data with a documented edit.

Why two kinds. TED indexes only the current version of a notice. When publication 201-2025 sits at version 3, the API cannot return version 1. A family where a correction genuinely changed the criteria therefore looks, through the API, like a single document with one set of criteria, because the superseded revision is gone. Testing whether the assessor applies a superseding document is impossible without supplying the superseded one.

How the difference is kept honest. Every case declares its source. A derived case carries a derivation field naming the real publications it came from and exactly what changed, and a derived case without that note fails validation. Labelling which cases were edited is what keeps those fixtures rather than fabrications.

The set is also mutation tested, because a suite that passes is not evidence that it would fail. Two tests break the assessor on purpose and assert the gate notices:

Deliberate bug What the gate reports
Use the earliest published criteria, ignoring corrections verdict 0.77, provenance 0.59, amendment recall 0.00
Never compare weights or names amendment recall 0.75
Ignore as_of entirely as_of isolation 0.92

The third row is why as_of isolation exists. Before it was added, ignoring as_of still passed the gate, because only four cases exercised it and four wrong answers out of 26 does not move an accuracy average.

10. Replay

What. Reconstructs one past decision and verifies it.

How. Nine steps, no cleverness. Load the run and the version it executed under. Verify the whole hash chain. Load the run's ledger steps. Re-check every citation against its stored document. Load the documents as of the run's date. Load the assessment from the ledger. Load the approver and their rationale. Load the scorecard. Load the trace.

Why write it first. It defines what every other part has to store. Building it last is how a project discovers in its final week that something was never recorded and cannot be recovered. This one was written in week two, before the agent existed.

If the chain check or the citation check fails, replay exits non-zero and names the row. Two tests deliberately break things and confirm it does.

11. The evidence bundle

What. A document rendering one run under the nine headings of Annex IV of Regulation (EU) 2024/1689.

How. replay --bundle out.html, or GET /api/runs/{run_id}/bundle, which the review page links to. Each heading is labelled with where its content came from:

Heading Source
1. General description of the system authored, run version data injected
2. Elements and the development process generated: TED scope, segmentation, retrieval, prompt hash, scorecard
3. Monitoring, functioning and control generated: ledger timeline, spans, the approval record
4. Appropriateness of the performance metrics authored, thresholds injected
5. Risk management system, Article 9 authored, and explicit about what the code does not do
6. Changes through the lifecycle generated from the distinct versions that have run
7. Harmonised standards applied out of scope, none claimed
8. EU declaration of conformity out of scope, cannot come from a repository
9. Post-market performance monitoring authored, including the retention position

Why label the sources. A generated placeholder in a legal document is worse than a visible gap. Sections 5 and 9 say plainly which obligations the software does not discharge, so whoever deploys it can see what they still owe.

Output is HTML only. Markdown was in the original plan and got cut, because it would have meant a second copy of three hundred lines of near-legal prose, and two copies of that drift. Browsers print to PDF.

12. The web app, and who gets to approve

What. A React client with three screens: the approval queue, the review page, and a list of runs.

How the review page earns its place. Each cited piece of evidence is rendered with the publication number, document kind, date and revision it came from:

"Award criterion 1: type price; name unstated; weight 80 (poi-exa)."
  666684-2023 · corrigendum · published 2023-11-02 · v1 · sha256 7d1c4f2a9b0e3861...

Why that detail. An approver signing off on evidence without seeing which version it came from defeats the point of storing the versions. It is the difference between a person reviewing a decision and a person rubber-stamping a screen.

Why the identity comes from the session. The approval request carries a decision and a rationale, and no field for an approver. The user is read from a signed server-side session and reloaded from the database on every request, so revoking someone's approver flag takes effect at once. A test posts a different approver_id in the body and confirms the ledger still records the session's user.

The command line can also record approvals, and there --as is a plain string: anyone holding database credentials can record a decision under any name. That is acceptable for an operator tool and not acceptable as the only path, which is why the web app exists.

13. Choosing a model

What. The component that reaches a verdict is a config string.

DECISION_LOG_ASSESSOR_MODEL=rules                     # default, free, deterministic
DECISION_LOG_ASSESSOR_MODEL=anthropic:claude-sonnet-5
DECISION_LOG_ASSESSOR_MODEL=openai:gpt-4.1
DECISION_LOG_ASSESSOR_MODEL=ollama:llama3.1           # fully local

How. init_chat_model ships inside LangChain, which LangGraph already pulls in, so there is no adapter layer, no factory and no plugin registry. Provider packages are optional extras: uv sync --extra anthropic.

Why it costs nothing in audit terms. Because replay reconstructs rather than re-runs, changing the model affects what future runs decide and has no effect on whether past runs can be reopened. The model id, prompt version and prompt hash are recorded with every run. Both the rules path and the model path return the same validated schema, so a model that invents an uncited finding is rejected before it reaches the record.

14. Retention

Records are kept by never deleting them. The application's database role holds no DELETE privilege on the ledger, the document revisions, the segments, the approvals or the scorecards, and there is no deletion path in the code. Meeting a retention floor therefore depends on the backup and recovery policy of the database, which belongs to whoever deploys this.

Why procurement

TED is free, public, needs no API key, and its documents genuinely get amended. Amendments are what make provenance hard, so a domain without them would not be a real test of any of the above.

Stack

Piece Why
Python 3.13 Typed throughout, checked by mypy in strict mode
PostgreSQL 18 Stores documents, the ledger, and graph checkpoints
LangGraph Only for interrupt(), which pauses a run for days and resumes it
FastAPI JSON API and session auth
React 19 + TypeScript + Vite The approval client
OpenTelemetry Spans written to local JSONL files
pydantic Validates every external input and every model output
Docker Compose Postgres and the app
uv Dependency management

No vector database, no ORM, no message queue, no Kubernetes. See What changes at scale.

Requirements

  • Docker and Docker Compose
  • uv
  • Node 20.19 or newer, only if you want to run the client's dev server. frontend/.nvmrc pins 24

No API key is needed. The default assessor is a deterministic rules engine, so a fresh clone runs the whole pipeline for free.

Quick start

git clone https://github.com/katiiab01/decision-log.git
cd decision-log

cp .env.example .env
# set MIGRATE_DB_PASSWORD, APP_DB_PASSWORD and DECISION_LOG_SESSION_SECRET

make up          # build and start Postgres 18 and the app
make migrate     # create tables, the two database roles, and the checkpoint tables
uv sync          # install python dependencies for the cli

Then reproduce the worked example from the top of this file:

uv run decision-log ingest --procedure 3978f4f3-4130-472a-831a-3eb7889b4470

# score the labelled set. a run cannot start without a passing scorecard
uv run decision-log eval

# create an approver. reads a piped password, or prompts if there is a terminal
printf 'a-long-password\n' | uv run decision-log user \
  --email you@example.org --name "Your Name" --approver

uv run decision-log run --procedure 3978f4f3-4130-472a-831a-3eb7889b4470 \
  --as-of 2024-06-01T00:00:00+00:00

uv run decision-log approve --run <run-id> --as you@example.org \
  --decision approved --rationale "why"

uv run decision-log replay <run-id>
uv run decision-log replay <run-id> --bundle bundle.html

Then open the app and do the same thing in a browser. make up publishes it on http://localhost:8000 by default, or set APP_PORT if that port is taken.

ingest without --procedure discovers procedures using the CPV code, countries and start date in your .env.

Working on the code

Run the API and the client separately so both reload:

make db          # Postgres only
make migrate
make api         # FastAPI on :8000
make web         # Vite on :5173, proxying /api to :8000

Open http://localhost:5173. The dev server proxies /api, so the browser only ever sees one origin. That is why there is no CORS configuration anywhere, and why DECISION_LOG_BASE_URL is http://localhost:5173 while you work this way.

Try breaking it

Connect to Postgres as the migration user, edit a stored document, and run replay again. It will fail and name the row:

docker compose exec db psql -U dlog_migrate -d decisionlog \
  -c "update chunk set text = 'quietly rewritten' where section = 'award_criterion'"

uv run decision-log replay <run-id>   # exits non-zero, reports the bad segment

Commands

Command Does
ingest Fetch TED notices and store immutable revisions
run Review a procedure and pause for approval
approve Record a decision and resume the run
replay Reconstruct a past decision and verify it. --bundle out.html writes the evidence bundle
verify-chain Check the whole ledger and its anchors
anchor Record the chain head outside the database
eval Score the labelled set and write a scorecard
user Create an account

Testing

make test              # pytest
make lint              # forbidden phrasing, ruff, mypy --strict
make build-frontend    # tsc --noEmit, then vite build

Tests need Postgres running (make db) and never touch the network. TED responses are recorded fixtures.

They run against a separate decisionlog_test database, created and migrated automatically on first run, because they truncate every table between tests. Your working data is left alone, and a guard refuses to run if the redirect ever fails.

The suite covers the claims that matter:

  • editing a ledger row, deleting one, or rewriting a payload is detected and named
  • the application's database role cannot UPDATE or DELETE the ledger
  • a run pauses, the connection and every cache are destroyed, and it still resumes
  • a document published after the run's as_of is invisible to it, and editing a cited document makes replay fail
  • a finding without a citation cannot be constructed, and a model returning one is rejected before it reaches the record
  • the evaluation gate notices an assessor that ignores corrections
  • an approval records the session's user even when the request body names another
  • an approver rationale containing markup is escaped in the rendered bundle

Project structure

migrations/            plain .sql files and a small runner
src/decision_log/
  config.py            settings from the environment
  db.py                connection pool
  hashing.py           canonical JSON and SHA-256
  ingest/              TED client, normalisation, segmentation, storage
  retrieval/           as-of filtered lookup, and full-text search
  agent/               the four graph steps, and the two assessors
  api/                 json endpoints, session auth, spa serving
  bundle/              the annex iv document, one jinja template
  ledger/              append, verify, and anchor the hash chain
  evaluation/          labelled set, scorer, scorecard
  telemetry/           OpenTelemetry spans to JSONL
  replay.py            reconstruct a past decision
  users.py             accounts and password hashing
  cli.py               entry point
frontend/              react client: login, queue, review
tests/

What this is not

The output is an evidence bundle mapped to Annex IV headings. It is not a claim that any system meets the AI Act, and nothing here attests to anything. No repository can, and saying otherwise would be a real liability rather than a marketing choice.

The ledger is tamper-evident, not tamper-proof. It detects changes made with the application's credentials, and changes made directly in the database without also editing the anchor file kept outside it. It does not stop a database superuser who rewrites the ledger and regenerates every hash after it. Keeping that file off-host raises the bar; nothing on a single host removes the limit.

The agent does not award contracts, does not rank tenders, and takes no action on its own. It prepares a finding and a person decides.

What changes at scale

Compose is enough to run and understand this. If it had to carry real load:

  • Ledger writes are serialised by an advisory lock. Shard the chain per run and anchor the run heads into a parent chain.
  • Retrieval is a structured lookup over small documents. Larger corpora would want the full-text index that is already generated on every segment, and embeddings after that.
  • Spans go to local files. Swap the exporter for OTLP, which is one line.
  • Postgres is a single container. Managed instance, read replicas for replay, point-in-time recovery for the retention floor.
  • The app is stateless apart from the pool, so it scales horizontally as is.
  • Login throttling is a per-process counter. It needs a shared store once there is more than one instance.

Status

Everything in the plan is built: ingestion, ledger, chain anchoring, retrieval, the agent and its approval interrupt, the evaluation gate, replay, the CLI, the JSON API, the approval client, and the Annex IV bundle.

One honest gap. The model assessor is tested against a stub chat model, which covers the wiring and the schema validation, but it has never been run against a live model. The default rules path is what the labelled set scores.

Licence

MIT

About

AI agent that reviews EU public procurement awards against their published criteria

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages