Skip to content

Document indexing the Postgres projections table - #132

Open
jonassvalin wants to merge 4 commits into
mainfrom
docs-projection-indexing
Open

jonassvalin wants to merge 4 commits into
mainfrom
docs-projection-indexing

Conversation

@jonassvalin

@jonassvalin jonassvalin commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds an "Indexing Projections in Postgres" section to the README. It explains which index expressions match the SQL the projection store renders, why partial expression indexes mislead the planner, the recommended alternatives, and how prepared statements interact with the current rendering. Docs only; no code changes. Partitioning the table by name is deliberately left out, because the library doesn't ship or test a partitioned schema.

This is the first of four PRs from the included plan. PRs 2–4 change how the Postgres converters render path keys, the projection name and jsonb values. Each can be accepted or declined on its own. This one describes behaviour that holds today, whatever happens to the others, and is meant to start the discussion.

Motivation

A downstream service had list queries with a ~700ms floor in production (2–3.8s locally). All its expression indexes were partial (WHERE name = '…'), and Postgres doesn't use a partial index's statistics when estimating selectivity. So every equality filter was estimated at a fixed 0.5%, and the planner walked the wrong index. Non-partial indexes that lead with name fixed it. Nothing in the library's docs says how to index the shared projections table, or that index expressions must match the rendered jsonb_extract_path(state, …) form exactly.

Changes

Key Changes

  • README section, after "Finalising State", covering:
    • a composite index example that leads with name, built CONCURRENTLY;
    • which filter operators render jsonb_extract_path and which render jsonb_extract_path_text;
    • the partial-index statistics trap, and when partial indexes are still fine;
    • the btree entry size limit, and GIN jsonb_path_ops for CONTAINS on larger values;
    • the prepared-statement caveat: path keys and name are bound today, so generic plans can't use these indexes. Workarounds are plan_cache_mode = force_custom_plan or a pool with prepare_threshold=None;
    • a caution about CREATE STATISTICS memory use during ANALYZE on large state documents.
  • Changelog fragment.

Sources for the Postgres behaviour

Breaking Changes

None.

Migration Guide

None needed.

How to Verify

Automated Verification

  • Docs build: mise run docs:build
  • Unit tests pass: mise run test:unit (1747)
  • Type checking passes: mise run types:check
  • Linting passes: mise run lint:check
  • Formatting is correct: mise run format:check
  • Integration and component tests weren't run locally; CI covers them. The change is docs only.

Manual Verification

Checked every claim against Postgres 16.3, with 200,000 rows split across two projection names:

  • Partial index only: an equality filter matching 60,000 rows was estimated at 504 (0.5%), and a range filter at a third. With a composite (name, expression) index, the estimate was 59,954.
  • Generic plans (EXPLAIN (GENERIC_PLAN)): with the path key bound, no expression index is used; with a literal key, the composite index is. With name bound, a partial index isn't used; with a literal name, it is.
  • _text and GIN indexes: a jsonb_extract_path_text index matches IS NULL, and a GIN jsonb_path_ops index matches @>.
  • btree size limit: indexing a key holding a large array makes inserts fail with index row size 7216 exceeds btree version 4 maximum 2704.
  • README examples: the SQL runs as written against sql/create_projections_table.sql plus sql/create_projections_indices.sql, and the Python snippet's imports resolve.

The ANALYZE memory caution comes from the downstream investigation and is stated without version specifics. I haven't verified it independently.

Checklist

  • I have read the contributing guidelines
  • I have added/updated tests for my changes (docs only)
  • I have updated documentation as needed
  • I have added a changelog fragment (if user-facing changes)
  • Breaking changes are documented with migration guidance

Related Issues

None. The implementation plan for all four PRs and its review are included under meta/plans/ and meta/reviews/plans/.

Add a README section on which index expressions match the store's
queries, why partial expression indexes leave the planner without
statistics, the recommended composite and partitioned alternatives, and
how prepared statements interact with the current rendering, plus a
changelog fragment.
The library doesn't ship or test a partitioned projections schema, so
recommending one is premature.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant