An agent that reviews EU public procurement awards against their published criteria, pauses for a human to sign off, and leaves a record you can reopen months later.
The agent is not the point. The record is.
Contents. The problem · A worked example · The one idea the rest depends on · How each part works · Stack · Quick start · Commands · Testing · What this is not · What changes at scale · Status
A city publishes a call for tenders. The notice says the winner is chosen on price, weighted 80, and quality, weighted 20.
Three weeks later the city publishes a correction. Nothing unusual, it happens constantly. Six months after that the contract is awarded, and a company that lost asks which criteria its bid was actually judged against.
If software helped review that file, you need an answer. And the documents have changed in the meantime, which is where it gets hard: did the software miss the correction, or had the correction not been published yet when it looked? Without a record of what it was looking at, you cannot tell the two apart.
That is the question this repository answers, for one decision at a time.
Everything below uses one real procurement, so you can follow the same steps and get
the same output. Procedure 3978f4f3-4130-472a-831a-3eb7889b4470 is a Finnish
catering contract. TED holds three publications for it.
$ decision-log ingest --procedure 3978f4f3-4130-472a-831a-3eb7889b4470
3978f4f3-4130-472a-831a-3eb7889b4470: 3 new revisions
Those three documents, oldest first:
| Published | Kind | Publication | Award criteria |
|---|---|---|---|
| 2023-10-13 | contract notice | 622805-2023 |
price 80, quality 20 |
| 2023-11-02 | corrigendum | 666684-2023 |
price 80, quality 20 |
| 2024-01-03 | award notice | 3869-2024 |
price 80 Kokonaishinta, quality 20 Laatu |
The middle row is the interesting one. It is a correction that restates the criteria,
so it replaces the original notice. Any answer about "the published criteria" has to
come from 666684-2023, not from 622805-2023.
Now review it, as the paperwork stood on 1 June 2024:
$ decision-log run --procedure 3978f4f3-4130-472a-831a-3eb7889b4470 \
--as-of 2024-06-01T00:00:00+00:00
run 108d72bd-89b7-4888-9ddb-d2b6163bfe31 is awaiting approval
It stopped. A person has to decide, and until they do, nothing is recorded as a decision. Marie approves it:
$ decision-log approve --run 108d72bd-89b7-4888-9ddb-d2b6163bfe31 \
--as marie@example.org --decision approved \
--rationale "The corrigendum restates price 80 / quality 20 and the award notice matches."
run 108d72bd-89b7-4888-9ddb-d2b6163bfe31 approved by marie@example.org
Months later, this is the payoff:
$ decision-log replay 108d72bd-89b7-4888-9ddb-d2b6163bfe31
run 108d72bd-89b7-4888-9ddb-d2b6163bfe31
as of 2024-06-01 00:00:00+00:00
agent 2296da5db77152dbdaea94e2943aebafc9f72d0a / rules / rules
chain 3 entries, intact
citations 6 verified, no errors
scorecard 1
spans 5
revisions as they stood:
2023-10-13 contract_notice 622805-2023 v1
2023-11-02 corrigendum 666684-2023 v1
2024-01-03 award_notice 3869-2024 v1
verdict consistent
summary 2 criteria compared, 0 mismatched, across 3 revisions.
approved by Marie Dubois <marie@example.org> on 2026-09-07: The corrigendum
restates price 80 / quality 20 and the award notice matches.
Every line is carrying something:
revisions as they stoodlists the documents that existed on 1 June 2024, not the documents that exist today.chain 3 entries, intactis a check that nobody edited the log afterwards.citations 6 verifiedmeans the six quoted pieces of evidence still match the documents, character for character.scorecard 1is the test run that was passing when the decision was made.approved byis a named person and their reasoning, with a timestamp.
decision-log replay <run-id> --bundle out.html writes the same thing as a document
laid out under the nine headings of Annex IV of the EU AI Act.
Replay reconstructs. It never re-runs the model.
The tempting design is: to find out what the AI decided, run the AI again on the same inputs. That is wrong twice over. Language models are not deterministic, so you would get a different answer and no way to tell a bug from ordinary variation. And even with a perfectly repeatable model, running it today tells you what today's system does, while the question was about that system on that day.
So every step writes its output into rows that cannot be edited, and replay is careful reading. No model is ever invoked during a replay.
What. Downloads procurement notices from TED, the EU tenders database, and stores every publication as a row that is never updated.
How. TED groups a procurement's paperwork under a procedure-identifier. All three
documents in the example above share one. The idempotence key is a hash of the response
content, so anything that changed becomes a new revision and an unchanged re-fetch does
nothing. Running ingest twice is safe.
Why content and not a version number. TED both republishes a notice under a new number and overwrites one in place, and it only indexes the current version. A hash makes no assumption that the source is honest or consistent about announcing its own changes. For an audit record, "store what we saw, when we saw it" is a stronger guarantee than "store what the source says changed."
TED's links field, a per-language block of PDF URLs, is excluded from that hash. It is
navigation rather than content, and hashing it would turn any change to TED's link
structure into a phantom revision of every notice.
What. Every run is pinned to a date and can only see documents published on or before it.
How. One SQL clause:
where notice_id = %s
and published_at <= %sWhy it matters. This is the whole of "the documents as they stood that day". No
snapshots, no time-travel database, one predicate over rows that never change. The
example above run with --as-of 2023-11-01 would see only the original notice and
report insufficient_evidence, because the award notice did not exist yet. Same code,
same database, different point in time.
What. Every claim the system makes points at a specific piece of a specific document.
How. A citation stores the revision, the segment, and the SHA-256 the text had at the time. The document is split into the smallest units worth quoting, each hashed:
Award criterion 1: type price; name unstated; weight 80 (poi-exa).
Why a pointer and not a copy. Copied text proves nothing, because anyone could have typed it. A pointer plus a hash is checkable: replay looks the segment up, re-hashes it, and either it matches or it fails loudly. A finding with no citation cannot even be constructed, because the schema requires at least one. An uncited claim in an audit record is worse than no record, since it looks like evidence and is not.
What. Four steps: gather the evidence, assess it, wait for a human, record the decision.
How. The waiting step calls LangGraph's interrupt(). The graph stops mid-function
and its state is written to Postgres. The process can be killed, the container rebuilt,
the machine rebooted. Days later someone approves, and execution resumes from inside
that function as if nothing happened.
Why a graph framework at all. Only for that. Writing it yourself means designing a state machine, a serialisation format and a resume protocol. Nothing else here needs LangGraph, and the README should say so rather than implying a broader dependency.
A test proves it: the connection is closed, the pool is closed, every cache is cleared, and only then does the run resume and complete.
What. An append-only log where each step of every run is recorded.
How. Each row stores a hash of its own contents and the hash of the row before it:
| seq | step | output hash | prev hash | entry hash |
|---|---|---|---|---|
| 1 | gather |
a3f2... |
0000... |
7d1c... |
| 2 | assess |
9b0e... |
7d1c... |
e4a8... |
| 3 | record |
1f77... |
e4a8... |
55b2... |
Why chain them. It makes the rows interdependent. Change row 1 and its hash
changes, so row 2's prev hash no longer matches, and so on down the file. You cannot
quietly edit one detail.
This is a hash chain: rows that commit to their predecessor. There is no network, no consensus and no distributed ledger involved, and it is worth being precise about that so nobody arrives expecting one.
One implementation note. seq is assigned by hand under an advisory lock rather than by
a database sequence, because the hash covers the row's position, so the position has to
be known before the insert and a sequence assigns it during.
What. The application connects with an account that cannot rewrite the record.
How. Eight lines of SQL:
grant select, insert on all tables in schema public to dlog_app;
grant update (status) on run to dlog_app;
revoke update, delete on notice_revision, chunk, ledger_entry,
approval, eval_scorecard from dlog_app;A separate account owns the schema and is used only by make migrate. The one exception
is run.status, because a run legitimately moves from running to awaiting_approval
to approved, and it is a column-level grant rather than a table-level one.
Why bother, when the code simply does not write those statements. Because "our code
does not do that" is a promise and "the credential cannot do that" is a fact. A bug, a
bad merge or an injection somewhere gets a permission error instead of a silent rewrite.
A test connects as the application role and confirms UPDATE on the ledger is refused.
What. decision-log anchor writes the chain's current tip to a file outside the
database.
How. It appends {at, seq, entry_hash} to data/anchors/heads.jsonl, and refuses
to anchor a chain that does not already verify. verify-chain then checks every
recorded tip against the ledger.
Why. Without it, someone with database superuser access could rewrite a row and regenerate every hash after it, and the chain would verify cleanly. With it, they would have to edit the database and the file. That is a higher bar, not an impossible one, and the honest limit of what a single host can offer.
An anchor file belongs to the ledger it anchors. Resetting a development database leaves
the anchors pointing at rows that no longer exist, and verify-chain will say so until
you remove them; make reset discards both together. In production that pairing is the
point: you cannot drop the database and quietly carry on.
What. A run cannot start unless the system's tests were passing.
How. decision-log eval scores a labelled set and writes one eval_scorecard row.
run.scorecard_id is a not null foreign key pointing at it.
Why a foreign key. Because it is then not a policy document, a checklist item or a CI step someone can skip. The database refuses the insert. Five properties are scored, and two of them have to be perfect:
| Metric | Threshold | What it means |
|---|---|---|
| Verdict accuracy | 0.80 | Did it reach the right conclusion |
| Provenance accuracy | 0.95 | Did it read the document it was supposed to read |
| Amendment recall | 0.90 | On cases where a correction changed the criteria, both of the above |
| Citation validity | 1.00 | Every quote resolves and still hashes to what was recorded |
as_of isolation |
1.00 | No run cited a document published after its own cut-off |
A model getting a hard judgement call wrong is a disclosed limitation. A model inventing a citation, or reading a document that did not exist yet, is a fabricated record. That is why the last two are invariants and not accuracy scores.
The gate also refuses a labelled set with fewer than eight amended cases, because a recall figure over one or two cases is a number rather than a measurement.
What. 26 cases, 16 taken verbatim from TED and 10 built from real data with a documented edit.
Why two kinds. TED indexes only the current version of a notice. When publication
201-2025 sits at version 3, the API cannot return version 1. A family where a
correction genuinely changed the criteria therefore looks, through the API, like a single
document with one set of criteria, because the superseded revision is gone. Testing
whether the assessor applies a superseding document is impossible without supplying the
superseded one.
How the difference is kept honest. Every case declares its source. A derived case
carries a derivation field naming the real publications it came from and exactly what
changed, and a derived case without that note fails validation. Labelling which cases
were edited is what keeps those fixtures rather than fabrications.
The set is also mutation tested, because a suite that passes is not evidence that it would fail. Two tests break the assessor on purpose and assert the gate notices:
| Deliberate bug | What the gate reports |
|---|---|
| Use the earliest published criteria, ignoring corrections | verdict 0.77, provenance 0.59, amendment recall 0.00 |
| Never compare weights or names | amendment recall 0.75 |
Ignore as_of entirely |
as_of isolation 0.92 |
The third row is why as_of isolation exists. Before it was added, ignoring as_of
still passed the gate, because only four cases exercised it and four wrong answers out
of 26 does not move an accuracy average.
What. Reconstructs one past decision and verifies it.
How. Nine steps, no cleverness. Load the run and the version it executed under. Verify the whole hash chain. Load the run's ledger steps. Re-check every citation against its stored document. Load the documents as of the run's date. Load the assessment from the ledger. Load the approver and their rationale. Load the scorecard. Load the trace.
Why write it first. It defines what every other part has to store. Building it last is how a project discovers in its final week that something was never recorded and cannot be recovered. This one was written in week two, before the agent existed.
If the chain check or the citation check fails, replay exits non-zero and names the row. Two tests deliberately break things and confirm it does.
What. A document rendering one run under the nine headings of Annex IV of Regulation (EU) 2024/1689.
How. replay --bundle out.html, or GET /api/runs/{run_id}/bundle, which the review
page links to. Each heading is labelled with where its content came from:
| Heading | Source |
|---|---|
| 1. General description of the system | authored, run version data injected |
| 2. Elements and the development process | generated: TED scope, segmentation, retrieval, prompt hash, scorecard |
| 3. Monitoring, functioning and control | generated: ledger timeline, spans, the approval record |
| 4. Appropriateness of the performance metrics | authored, thresholds injected |
| 5. Risk management system, Article 9 | authored, and explicit about what the code does not do |
| 6. Changes through the lifecycle | generated from the distinct versions that have run |
| 7. Harmonised standards applied | out of scope, none claimed |
| 8. EU declaration of conformity | out of scope, cannot come from a repository |
| 9. Post-market performance monitoring | authored, including the retention position |
Why label the sources. A generated placeholder in a legal document is worse than a visible gap. Sections 5 and 9 say plainly which obligations the software does not discharge, so whoever deploys it can see what they still owe.
Output is HTML only. Markdown was in the original plan and got cut, because it would have meant a second copy of three hundred lines of near-legal prose, and two copies of that drift. Browsers print to PDF.
What. A React client with three screens: the approval queue, the review page, and a list of runs.
How the review page earns its place. Each cited piece of evidence is rendered with the publication number, document kind, date and revision it came from:
"Award criterion 1: type price; name unstated; weight 80 (poi-exa)."
666684-2023 · corrigendum · published 2023-11-02 · v1 · sha256 7d1c4f2a9b0e3861...
Why that detail. An approver signing off on evidence without seeing which version it came from defeats the point of storing the versions. It is the difference between a person reviewing a decision and a person rubber-stamping a screen.
Why the identity comes from the session. The approval request carries a decision and
a rationale, and no field for an approver. The user is read from a signed server-side
session and reloaded from the database on every request, so revoking someone's approver
flag takes effect at once. A test posts a different approver_id in the body and
confirms the ledger still records the session's user.
The command line can also record approvals, and there --as is a plain string: anyone
holding database credentials can record a decision under any name. That is acceptable
for an operator tool and not acceptable as the only path, which is why the web app
exists.
What. The component that reaches a verdict is a config string.
DECISION_LOG_ASSESSOR_MODEL=rules # default, free, deterministic
DECISION_LOG_ASSESSOR_MODEL=anthropic:claude-sonnet-5
DECISION_LOG_ASSESSOR_MODEL=openai:gpt-4.1
DECISION_LOG_ASSESSOR_MODEL=ollama:llama3.1 # fully local
How. init_chat_model ships inside LangChain, which LangGraph already pulls in, so
there is no adapter layer, no factory and no plugin registry. Provider packages are
optional extras: uv sync --extra anthropic.
Why it costs nothing in audit terms. Because replay reconstructs rather than re-runs, changing the model affects what future runs decide and has no effect on whether past runs can be reopened. The model id, prompt version and prompt hash are recorded with every run. Both the rules path and the model path return the same validated schema, so a model that invents an uncited finding is rejected before it reaches the record.
Records are kept by never deleting them. The application's database role holds no
DELETE privilege on the ledger, the document revisions, the segments, the approvals or
the scorecards, and there is no deletion path in the code. Meeting a retention floor
therefore depends on the backup and recovery policy of the database, which belongs to
whoever deploys this.
TED is free, public, needs no API key, and its documents genuinely get amended. Amendments are what make provenance hard, so a domain without them would not be a real test of any of the above.
| Piece | Why |
|---|---|
| Python 3.13 | Typed throughout, checked by mypy in strict mode |
| PostgreSQL 18 | Stores documents, the ledger, and graph checkpoints |
| LangGraph | Only for interrupt(), which pauses a run for days and resumes it |
| FastAPI | JSON API and session auth |
| React 19 + TypeScript + Vite | The approval client |
| OpenTelemetry | Spans written to local JSONL files |
| pydantic | Validates every external input and every model output |
| Docker Compose | Postgres and the app |
| uv | Dependency management |
No vector database, no ORM, no message queue, no Kubernetes. See What changes at scale.
- Docker and Docker Compose
- uv
- Node 20.19 or newer, only if you want to run the client's dev server.
frontend/.nvmrcpins 24
No API key is needed. The default assessor is a deterministic rules engine, so a fresh clone runs the whole pipeline for free.
git clone https://github.com/katiiab01/decision-log.git
cd decision-log
cp .env.example .env
# set MIGRATE_DB_PASSWORD, APP_DB_PASSWORD and DECISION_LOG_SESSION_SECRET
make up # build and start Postgres 18 and the app
make migrate # create tables, the two database roles, and the checkpoint tables
uv sync # install python dependencies for the cliThen reproduce the worked example from the top of this file:
uv run decision-log ingest --procedure 3978f4f3-4130-472a-831a-3eb7889b4470
# score the labelled set. a run cannot start without a passing scorecard
uv run decision-log eval
# create an approver. reads a piped password, or prompts if there is a terminal
printf 'a-long-password\n' | uv run decision-log user \
--email you@example.org --name "Your Name" --approver
uv run decision-log run --procedure 3978f4f3-4130-472a-831a-3eb7889b4470 \
--as-of 2024-06-01T00:00:00+00:00
uv run decision-log approve --run <run-id> --as you@example.org \
--decision approved --rationale "why"
uv run decision-log replay <run-id>
uv run decision-log replay <run-id> --bundle bundle.htmlThen open the app and do the same thing in a browser. make up publishes it on
http://localhost:8000 by default, or set APP_PORT if that port is taken.
ingest without --procedure discovers procedures using the CPV code, countries and
start date in your .env.
Run the API and the client separately so both reload:
make db # Postgres only
make migrate
make api # FastAPI on :8000
make web # Vite on :5173, proxying /api to :8000Open http://localhost:5173. The dev server proxies /api, so the browser only ever sees
one origin. That is why there is no CORS configuration anywhere, and why
DECISION_LOG_BASE_URL is http://localhost:5173 while you work this way.
Connect to Postgres as the migration user, edit a stored document, and run replay
again. It will fail and name the row:
docker compose exec db psql -U dlog_migrate -d decisionlog \
-c "update chunk set text = 'quietly rewritten' where section = 'award_criterion'"
uv run decision-log replay <run-id> # exits non-zero, reports the bad segment| Command | Does |
|---|---|
ingest |
Fetch TED notices and store immutable revisions |
run |
Review a procedure and pause for approval |
approve |
Record a decision and resume the run |
replay |
Reconstruct a past decision and verify it. --bundle out.html writes the evidence bundle |
verify-chain |
Check the whole ledger and its anchors |
anchor |
Record the chain head outside the database |
eval |
Score the labelled set and write a scorecard |
user |
Create an account |
make test # pytest
make lint # forbidden phrasing, ruff, mypy --strict
make build-frontend # tsc --noEmit, then vite buildTests need Postgres running (make db) and never touch the network. TED responses are
recorded fixtures.
They run against a separate decisionlog_test database, created and migrated
automatically on first run, because they truncate every table between tests. Your
working data is left alone, and a guard refuses to run if the redirect ever fails.
The suite covers the claims that matter:
- editing a ledger row, deleting one, or rewriting a payload is detected and named
- the application's database role cannot
UPDATEorDELETEthe ledger - a run pauses, the connection and every cache are destroyed, and it still resumes
- a document published after the run's
as_ofis invisible to it, and editing a cited document makes replay fail - a finding without a citation cannot be constructed, and a model returning one is rejected before it reaches the record
- the evaluation gate notices an assessor that ignores corrections
- an approval records the session's user even when the request body names another
- an approver rationale containing markup is escaped in the rendered bundle
migrations/ plain .sql files and a small runner
src/decision_log/
config.py settings from the environment
db.py connection pool
hashing.py canonical JSON and SHA-256
ingest/ TED client, normalisation, segmentation, storage
retrieval/ as-of filtered lookup, and full-text search
agent/ the four graph steps, and the two assessors
api/ json endpoints, session auth, spa serving
bundle/ the annex iv document, one jinja template
ledger/ append, verify, and anchor the hash chain
evaluation/ labelled set, scorer, scorecard
telemetry/ OpenTelemetry spans to JSONL
replay.py reconstruct a past decision
users.py accounts and password hashing
cli.py entry point
frontend/ react client: login, queue, review
tests/
The output is an evidence bundle mapped to Annex IV headings. It is not a claim that any system meets the AI Act, and nothing here attests to anything. No repository can, and saying otherwise would be a real liability rather than a marketing choice.
The ledger is tamper-evident, not tamper-proof. It detects changes made with the application's credentials, and changes made directly in the database without also editing the anchor file kept outside it. It does not stop a database superuser who rewrites the ledger and regenerates every hash after it. Keeping that file off-host raises the bar; nothing on a single host removes the limit.
The agent does not award contracts, does not rank tenders, and takes no action on its own. It prepares a finding and a person decides.
Compose is enough to run and understand this. If it had to carry real load:
- Ledger writes are serialised by an advisory lock. Shard the chain per run and anchor the run heads into a parent chain.
- Retrieval is a structured lookup over small documents. Larger corpora would want the full-text index that is already generated on every segment, and embeddings after that.
- Spans go to local files. Swap the exporter for OTLP, which is one line.
- Postgres is a single container. Managed instance, read replicas for replay, point-in-time recovery for the retention floor.
- The app is stateless apart from the pool, so it scales horizontally as is.
- Login throttling is a per-process counter. It needs a shared store once there is more than one instance.
Everything in the plan is built: ingestion, ledger, chain anchoring, retrieval, the agent and its approval interrupt, the evaluation gate, replay, the CLI, the JSON API, the approval client, and the Annex IV bundle.
One honest gap. The model assessor is tested against a stub chat model, which covers the
wiring and the schema validation, but it has never been run against a live model. The
default rules path is what the labelled set scores.
MIT