Skip to content

Name event columns explicitly in Postgres event storage adapter - #131

Open
wickedwukong wants to merge 13 commits into
mainfrom
gh-130-explicit-event-columns
Open

wickedwukong wants to merge 13 commits into
mainfrom
gh-130-explicit-event-columns

Conversation

@wickedwukong

Copy link
Copy Markdown

Summary

PostgresEventStorageAdapter now names the StoredEvent columns in every query whose rows reach class_row(StoredEvent), instead of using *. An events table with a column the installed StoredEvent does not know no longer breaks reads or writes.

Closes #130.

Motivation

psycopg's class_row calls StoredEvent(**row), so any extra column becomes an unexpected keyword argument:

TypeError: StoredEvent.__init__() got an unexpected keyword argument 'metadata'

This turns additive schema changes into read-time and write-time breaks. It forced a read shim onto both sides of a downstream 0.1.11 → 0.1.12a3 upgrade.

The issue names the SELECT * reads. Two more wildcards fed class_row and are fixed here too: RETURNING * (every save) and SELECT DISTINCT ON (...) * (save to a category).

Changes

Key Changes

Path Query builder Before
every save insert_batch_query RETURNING *
latest, save to an existing stream read_last_query SELECT *
scan scan_query Query().select_all()
save to a category with existing streams read_last_category_batch_query SELECT DISTINCT ON (category, stream ) *
  • The column list comes from one place: EVENT_COLUMNS, derived from dataclasses.fields(StoredEvent), and its SQL form EVENT_COLUMNS_SQL. class_row already relies on field names matching column names, so this adds no new coupling.
  • New changelog fragment for this fix and its limits.
  • Corrected the existing metadata changelog fragment (see below).

Implementation Details

  • scan_query uses Query().select(*EVENT_COLUMNS), so the constraint appliers (which only add WHERE/ORDER BY) are unaffected.
  • Star, select_all() and the generic query converters are unchanged. Their other users never build StoredEvent.
  • The projection, subscriber and subscription stores already use dict_row and pick columns by name, so they are unchanged.
  • Deliberately not done: a row factory that ignores unknown columns. It would silently drop data and hide column typos.

Changelog correction for metadata

The metadata fragment said the previous version "ignores the extra column", so an application-only rollback could leave it in place. That is not true: 0.1.11 reads with SELECT * and RETURNING * and fails as soon as the column exists. The fragment now:

  • splits the migration into two steps, keeping DEFAULT 'null'::jsonb until every writer supplies metadata;
  • states that old and new versions cannot run side by side against the migrated table, and that this fix cannot change that retroactively;
  • keeps the data-loss warning for DROP COLUMN.

Breaking Changes

None. No API or schema change. Queries now fetch only known columns, so an extra column is no longer read.

What this does and does not guarantee

  • From this release on, the running version tolerates an extra column. That makes additive migrations safe during a rolling deploy, provided a new NOT NULL column keeps a default while old and new versions overlap, and the migration runs before the new version starts.
  • It does not make the 0.1.11 → 0.1.12 metadata upgrade rolling-safe. Pods still on 0.1.11 break once metadata exists. A 0.1.11.x backport of this fix, deployed first, would make that upgrade safe.
  • Non-additive changes (renames, drops, type changes) are still out of scope.

How to Verify

  • mise run --force: 1747 unit, 143 integration, 3 component tests pass; lint, format and types clean.
  • TestPostgresStorageAdapterUnknownColumns adds unknown_column TEXT NOT NULL DEFAULT 'unknown' to a real table and covers every class_row path: save to a new stream, save to an existing stream, latest (log, category, stream), scan (log, category, stream), and save to a category with existing streams. Each test was seen failing with the TypeError before its query change.
  • test_batch_insert_query_builds_correct_sql now pins the explicit RETURNING list.

The plan and research for this change are in meta/plans/ and meta/research/.

🤖 Generated with Claude Code

insert_batch_query returned *, so an events table with a column
StoredEvent does not know broke every save with TypeError. Derive
the column list from StoredEvent's fields and use it in RETURNING.

Part of #130.
read_last_query selected *, so an events table with a column
StoredEvent does not know broke latest and every save to an
existing stream with TypeError. Select the StoredEvent columns
by name instead.

Part of #130.
scan_query selected *, so an events table with a column
StoredEvent does not know broke scan with TypeError. Select the
StoredEvent columns by name instead.

Part of #130.
read_last_category_batch_query selected DISTINCT ON ... *, so an
events table with a column StoredEvent does not know broke every
save to a category with existing streams with TypeError. Select
the StoredEvent columns by name instead.

Part of #130.
event_columns() took no input and always built the same SQL, so
hold it as EVENT_COLUMNS_SQL instead.

Part of #130.
Add a fragment for the explicit event columns fix and its limits.

Correct the metadata fragment: releases before metadata support,
such as 0.1.11, do not ignore the extra column. Their wildcard
reads and RETURNING fail once it exists, so an application-only
rollback does not work. Split the migration so the default stays
while any writer omits metadata.

Part of #130.
@wickedwukong
wickedwukong force-pushed the gh-130-explicit-event-columns branch from d6b7587 to dbf70c1 Compare October 5, 2026 09:43

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Postgres adapter reads break when the events table has columns the installed StoredEvent does not know about

1 participant