Skip to content

feat(server): index the agent catalog from the identity registry (#17) - #118

Merged
moises-cisneros merged 5 commits into
mainfrom
feat/17-agent-catalog-index
Oct 9, 2026
Merged

moises-cisneros merged 5 commits into
mainfrom
feat/17-agent-catalog-index

Conversation

@fercodes

@fercodes fercodes commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Refs #17

Summary

The catalog endpoint read the whole catalog from the chain on every request and
only cached it in memory. This adds a durable Postgres index synced from the
identity registry, with a resume cursor, and serves list/get from it while
keeping the current endpoint surface unchanged.

  • Ledger: getLatestLedger and getEvents (xdrFormat: json) on
    SorobanRpcClient (ledger range xor cursor); registry_events.dart decodes
    registered, metadata_set, uri_updated; RegistryEventReader port +
    SorobanRegistryEventReader (cursor pagination).
  • Index: AgentRecord + CatalogIndexState models and migration;
    AgentIndexRepository + ServerpodAgentIndexRepository (idempotent upsert by
    registry id, per-network resume cursor, newest-registration dedupe).
  • Indexer: CatalogIndexer bootstraps from registry state (the RPC keeps
    only ~7 days of events), then reads events after the cursor and re-hydrates
    the affected agents from state; any agent missing from the index is always
    recovered. A chain outage aborts the pass and leaves the index intact; a
    retention gap resets the cursor and recovers from state.
  • Read path: CatalogReader port, IndexedAgentCatalogService (index
    first, on-chain fallback while empty), agent_catalog_wiring.dart;
    AgentEndpoint depends only on CatalogReader.
  • Loop: CatalogIndexerLoopConfig + startCatalogIndexer, started from
    server.dart behind PULS3_INDEXER_ENABLED (off by default).
  • Docs: PULS3_INDEXER_* in .env.example, server README, api.md note,
    and openspec change artifacts under openspec/changes/17-agent-catalog-index/.

Acceptance criteria

  • After seeding, the list returns every demo agent — live testnet smoke:
    8 agents (registry ids 7–14), 7 orphans skipped. Output in evidence below.
  • Running the indexer twice creates no duplicates, and it resumes from the
    last processed ledger. Both are tested (unit + integration).
  • If RPC is down, the endpoint still serves from the index, and the indexer
    logs the error and retries on the next run. Tested.
  • Endpoint errors match docs: API contract between Flutter app and Serverpod #8 (AgentCatalogUnavailable unchanged).
  • Endpoints do not import the RPC client directly: AgentEndpoint now
    depends only on CatalogReader.
  • Search by skill returns only matching agents — deferred (see Notes).
  • Pagination works — deferred (see Notes).

Verification evidence

Run in puls3_server (local Postgres matching config/test.yaml):

$ dart analyze --fatal-infos
No issues found!
$ dart test test/unit
320 tests passed
$ dart test -j 1 test/integration
agent_index_repository_test, agent_endpoint_index_test, hire_repository_test,
chain_submission_store_test, chain_submission_tracker_restart_test,
greeting_endpoint_test, health_endpoint_test -> all passed
$ serverpod generate && git diff --exit-code
no diff (reproducible)
$ serverpod create-migration
"Server migration skipped. No changes detected."

Live testnet smoke run of CatalogIndexer against a real RPC (in-memory
repository, no DB):

catalog pass: from bootstrap to 5081240, events 0, touched 15, hydrated 8, skipped 7
catalog (8 agents):
#7 agt-001 | Ledger Scout | on-chain-analytics, monitoring, summaries | 5000000 stroops
#8 agt-002 | Remit Pilot | payments, anchors, compliance | 12500000 stroops
#9 agt-003 | Soroban Auditor | smart-contracts, security, rust | 45000000 stroops
#10 agt-004 | Invoice Clerk | document-parsing, payments, accounting | 2500000 stroops
#11 agt-005 | Market Pulse | on-chain-analytics, trading, summaries | 7500000 stroops
#12 agt-006 | Copy Forge | copywriting, marketing, summaries | 3000000 stroops
#13 agt-007 | Support Relay | customer-support, triage, monitoring | 1000000 stroops
#14 agt-008 | Data Weaver | data-cleaning, accounting, document-parsing | 6000000 stroops

Recorded testnet fixtures are committed for the RPC tests:
test/unit/ledger/fixtures/get_latest_ledger.json and
.../get_events_registry_register_7.json.

Notes for reviewers

  • Scope decision (agreed with maintainer): keep the current endpoint
    surface. Pagination and server-side search are deferred: the merged docs: API contract between Flutter app and Serverpod #8
    contract (CatalogEndpoint.listAgents unpaginated; F2-3/F2-4 filter
    client-side) does not define them, so feat: agent catalog endpoints with on-chain indexing #17's related ACs should be amended
    rather than this PR changing the app. docs/architecture/api.md records the
    deferral.
  • Model choice: a server-side AgentRecord (nullable catalog fields) rather
    than the domain AgentRepository, because the registry allows a registered
    agent with no wallet yet, and the domain Agent requires a wallet plus has no
    model/registryId.
  • The migration records the bundled Serverpod 4.0.4 module versions; the
    previous migration predates them. create-migration reports no pending
    changes, so the chain is consistent.
  • The indexer is off by default; set PULS3_INDEXER_ENABLED=true to run it.

The catalog endpoint read the whole catalog from the chain on every request
and only cached it in memory. Add a durable Postgres index synced from the
identity registry, with a resume cursor, and serve list/get from it.

- Ledger: `getLatestLedger` and `getEvents` (`xdrFormat: json`) on
  `SorobanRpcClient` (ledger range xor cursor); `registry_events.dart` decodes
  `registered`, `metadata_set` and `uri_updated`; `RegistryEventReader` port +
  `SorobanRegistryEventReader` (cursor pagination).
- Index: `AgentRecord` and `CatalogIndexState` models + migration,
  `AgentIndexRepository` port and `ServerpodAgentIndexRepository` (idempotent
  upsert by registry id, per-network resume cursor).
- Indexer: `CatalogIndexer` bootstraps from registry state (RPC keeps only
  ~7 days of events), then reads events after the cursor and hydrates the
  affected agents from state; any agent missing from the index is recovered.
  Chain outages abort the pass and leave the index intact; a retention gap
  resets the cursor and recovers from state.
- Read path: `CatalogReader` port, `IndexedAgentCatalogService` (index first,
  on-chain fallback while empty), `agent_catalog_wiring.dart`; `AgentEndpoint`
  depends only on `CatalogReader`.
- Loop: `CatalogIndexerLoopConfig` + `startCatalogIndexer`, started from
  `server.dart` behind `PULS3_INDEXER_ENABLED` (off by default).
- Tests: unit tests for the RPC events, parser, indexer and indexed service;
  integration tests for the repository and the endpoint. Recorded testnet
  fixtures for `getLatestLedger` and a registry `getEvents` page.
- Docs: `PULS3_INDEXER_*` in `.env.example`, server README, api.md note, and
  openspec change artifacts.

Refs #17

@TOMOKI977 TOMOKI977 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks Fernando, the overall design is solid. The cursor is written only after the upsert, bootstrap runs from registry state, events are filtered by the registry contract id, the models are serverOnly, and all writes go through the ORM. Requesting changes for a few edge cases and one regression.

Must fix

  1. The catalog cache regressed. buildAgentCatalogReader (agent_catalog_wiring.dart:48) runs on every request and builds a new AgentCatalogService(ledger) each time. The 60s cache, the shared in-flight refresh and serve-last-list-on-outage were all instance state, so they are gone on the fallback path. PULS3_INDEXER_ENABLED defaults to false, which means the default deployment reads the whole catalog from the chain on every list/get. The test that guarded "the default service is reused" was removed. Please build the fallback once per process (e.g. a lazily created singleton) and restore that test.

  2. get breaks the index-first contract. IndexedAgentCatalogService.get (line 30) falls back to the chain on any index miss, even when the index has rows. The indexed-agent-catalog spec and the README say reads must not touch the chain while the index has rows. Combined with point 1, get('unknown-id') triggers a full uncached chain read, and during an outage it throws instead of returning null. The test get serves the index, then the fallback asserts the opposite of the spec. Please only fall back when the index is empty and update that test.

  3. A gap in the event stream leaves stale rows forever. In catalog_indexer.dart:84-93, any RpcRequestRejected (that is, -32600/-32602, not only an out-of-retention start ledger) advances the cursor to latest. Only agents missing from the index are recovered. An already-indexed agent whose price, wallet or name changed inside the gap keeps serving the old values, because nothing ever re-reads it. SorobanRegistryEventReader also stops at _maxPages without signaling truncation (soroban_registry_event_reader.dart:45), so the cursor advances the same way. Suggested fix: on a gap, or on truncation, re-hydrate every indexed id (not just the missing ones), and make truncation an explicit result or exception. Please add a test where the index already has the agent and its metadata changes during the gap.

  4. The outage test does not await. In catalog_indexer_test.dart:246-255, expect(indexer(...).pass(), throwsA(...)) is not awaited. The "index untouched / checkpoint null" assertions therefore run before the pass executes and prove nothing. Use await expectLater(...) in an async test.

Should fix (non-blocking)

  • Orphan or invalid registrations never reach the index, so every pass re-reads them over RPC (about 7 sequential calls each, every 30s). Consider persisting skipped ids, or keeping a tombstone row that the read path excludes.
  • When an indexed agent's metadata becomes invalid, it is counted as skipped, but its old row keeps being served. The repository needs a delete or invalidate path.
  • AgentRecord is unique by registryId only. If PULS3_STELLAR_IDENTITY_REGISTRY changes, old rows count as already indexed. Scope rows by network and registry contract, like the cursor.
  • There is no single-runner guard. With multiple instances, the select-then-insert in upsertAll can hit unique violations. Document "one indexer instance", or use an upsert with ON CONFLICT.
  • There are no tests for SorobanRegistryEventReader (pagination, cursor, bound), for either wiring module, or for the LedgerException path where the cursor stays put.
  • Layering: agent/registry_event_reader.dart imports ledger/registry_events.dart, while the adapter in ledger/ imports from agent/. Moving the event types next to the port (or into a neutral module) removes the cycle.
  • Docs: api.md:26 says CatalogUnavailable where it should say AgentCatalogUnavailable. api.md:77 and api.md:343 still describe the in-memory cache and a "planned in #17" stale warning. The proposal lists agent-catalog-endpoint as modified but has no delta spec for it. Gates 6.1-6.3 in tasks.md are unchecked.
  • Nits: _rpcTimeout is now defined in three places, there are unused event fields (owner, uri, key, value, txHash), and some lines exceed 80 columns.

Process

  • Size: about 2.3k authored lines, against the ~700 in the forecast. Your tasks already suggest a clean 3-way split (RPC events → index + indexer → wiring). Splitting would make the re-review much faster.
  • The branch is based on c328902 and is behind main (#117 landed). After the rebase, please check that the migration timestamps still sort after the ones on main, and regenerate if needed.

Happy to pair on point 3 if useful.

…-index

# Conflicts:
#	puls3_server/migrations/migration_registry.txt
Address the review on #118 and update the branch with main.

- Cache regression: the on-chain fallback is now a process-wide singleton, so
  its 60 s cache, shared in-flight refresh and serve-last-list-on-outage survive
  across requests; a wiring test asserts it is built once.
- Index-first `get`: it falls back to the chain only when the index is empty, so
  an unknown id on a populated index returns `null` without reading the chain
  (test updated to the spec).
- Event-stream gaps: `RegistryEventReader.eventsSince` returns a
  `RegistryEventBatch` whose `truncated` flag makes a page-bound stop explicit;
  on a retention gap or truncation the indexer re-hydrates every indexed agent
  (not only the missing ones), so an agent whose metadata changed in the gap is
  refreshed. Tests cover both.
- The outage test now awaits the pass (`expectLater`).
- Moved the registry event types next to the port, removing the agent→ledger
  import cycle (`ledger/registry_events.dart` re-exports them).
- Docs: `api.md` says `AgentCatalogUnavailable` and describes the index-first
  outage behavior; added a delta spec for `agent-catalog-endpoint`; checked the
  gate boxes in `tasks.md`.
- Updated with main and regenerated the migration after the escrow-relay one.

Refs #17
@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Deploying puls3 with  Cloudflare Pages  Cloudflare Pages

Latest commit: e7e4e5e
Status:⚡️  Build in progress...

View logs

@fercodes

fercodes commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Thanks! Addressed the review in 72c0792 (updated with main):
Must fix

  1. Cache regression: the on-chain fallback is now a process-wide lazy singleton (chainCatalog in agent_catalog_wiring.dart), so the 60 s cache, the shared in-flight refresh and serve-last-list-on-outage survive across requests. Added agent_catalog_wiring_test.dart asserting it is built once (and that an injected environment stays isolated).
  2. Index-first get: it falls back only when the index is empty; an unknown id on a populated index returns null without touching the chain. The IndexedAgentCatalogService.get test is updated to the spec.
  3. Gap/truncation: RegistryEventReader.eventsSince now returns a RegistryEventBatch with an explicit truncated flag, set by the adapter when it stops at its page bound. On a retention gap (RpcRequestRejected) or truncation, CatalogIndexer re-hydrates every indexed id, not just the missing ones. Tests: an indexed agent whose name changed during a retention gap is refreshed; same for a truncated batch.
  4. Awaited outage test: now await expectLater(...) in an async test.
    Also
  • Moved the registry event types next to the port (agent/registry_events.dart); ledger/registry_events.dart re-exports them, removing the agent→ledger cycle.
  • Docs: api.md uses AgentCatalogUnavailable and describes the index-first outage behavior; added a delta spec for agent-catalog-endpoint; checked the tasks.md gate boxes.
  • Updated with main; regenerated the migration on top of the escrow-relay migration 20261008045629402 so timestamps sort.
    Deferred as follow-ups (non-blocking): persisting skipped/tombstoned ids so orphans are not re-read every pass; invalidating an indexed row whose metadata becomes invalid; scoping AgentRecord by network/registry; a single-runner guard / ON CONFLICT upsert; and splitting this PR into the 3 units. Happy to open issues for these.
    Merged main rather than rebasing — say the word if you prefer a rebase.

@fercodes
fercodes requested a review from TOMOKI977 October 9, 2026 18:45
@TOMOKI977

Copy link
Copy Markdown
Contributor

Thanks for addressing the earlier review. The cache, index-first lookup, gap recovery, and awaited outage test are fixed. Two blockers remain:

  1. When an affected agent becomes invalid, the indexer increments skipped but leaves the previous valid row in Postgres, then advances the cursor. Please add a delete/invalidate operation, remove invalid touched ids before checkpointing, and test valid -> invalid -> absent from list/get.
  2. Rows and checkpoints are not scoped by identity-registry contract. Changing PULS3_STELLAR_IDENTITY_REGISTRY on the same network can serve old rows and reuse the old cursor. Please scope both tables and repository queries by network + registry contract, use a composite unique key, pass the contract through both wiring paths, and add an isolation test.

After those fixes, update the branch from main and rerun CI.

@moises-cisneros moises-cisneros left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Status after resolving the merge conflicts (56a5494).

Conflicts

  • server.dart: kept both the catalog indexer and the wallet auth imports.
  • Generated protocol code and migrations: took main's version, then regenerated with serverpod generate. The old migration 20261009170004821 was computed without main's wallet tables, so I replaced it with 20261009213303097, which creates only agent_record and catalog_index_state.
  • The PR is MERGEABLE. dart analyze is clean and the 483 server unit tests pass locally.

Earlier review (must-fix items)
Checked against the code on the current head:

  1. The on-chain fallback is a process-wide singleton, built once.
  2. get falls back to the chain only when the index is empty.
  3. A gap or a truncated event page re-hydrates every indexed agent.
  4. The outage test uses await expectLater.

The non-blocking items and the size/split suggestion are not covered here.

CI is red for infrastructure reasons, not for this diff

  • server: fails at "Initialize containers" with toomanyrequests (Docker Hub unauthenticated pull rate limit) while pulling pgvector/pgvector:pg16. The tests never ran.
  • docker: 504 Gateway Timeout from auth.docker.io while resolving dart and alpine.
  • gate fails only because it requires those two.
  • feat/20-hire-runner fails the same way at the same time. Jobs that do not touch Docker (app, scripts, changes, scan) pass.
  • Failed jobs were re-run three times with the same result.

Next steps

  • Re-run the failed jobs once Docker Hub recovers, and confirm server and gate go green.
  • Longer term, authenticate the pulls (credentials: on the services and docker/login-action) with a Docker Hub token stored as a repo secret. That belongs in a separate CI PR.
  • @TOMOKI977 a re-review of the four items above would unblock this once CI is green.

@moises-cisneros
moises-cisneros merged commit 63f15e3 into main Oct 9, 2026
9 of 10 checks passed
@moises-cisneros
moises-cisneros deleted the feat/17-agent-catalog-index branch October 9, 2026 22:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants