Skip to content

feat(server): chain submission tracker for create_job and fund - #102

Merged
moises-cisneros merged 14 commits into
feat/97-chain-submission-trackerfrom
feat/97-tracker-loop
Oct 6, 2026
Merged

moises-cisneros merged 14 commits into
feat/97-chain-submission-trackerfrom
feat/97-tracker-loop

Conversation

@TOMOKI977

Copy link
Copy Markdown
Contributor

Closes #97 (scope for the Stellar Elite demo: create_job and fund confirmation)

Stacked on #100. Review and merge #100 first; this PR then retargets to main.

⚠️ Do not set PULS3_TRACKER_ENABLED=true until #96 lands. Until then the tracker uses NoopEscrowEffects, which confirms any successful create_job / fund that emitted an escrow event without verifying it against a hire. The flag is false by default.

Summary

The durable tracker from the relay design (docs/architecture/api.md, steps 4 and 5):

  • ChainSubmissionTracker.pass(), per submitted record:
    • NOT_FOUND: if the chain's close time is past validUntil, it fails with PreparationExpired; otherwise it resends the persisted envelope every resendAfter (30 s). A rejected envelope fails with SubmissionRejected.
    • FAILED: fails with TransactionFailed.
    • SUCCESS for createJob / fund: needs the matching escrow event (else JobEvidenceUnavailable), then calls EscrowEffects. EffectOk confirms the record; a failed effect fails it with JobMismatch / JobEvidenceUnavailable.
    • LedgerUnavailable or an exception leaves the record submitted for the next pass.
  • Escrow events: job_created and job_funded parsed from getTransaction, filtered by the configured escrow contract id. Fixtures are the real testnet create_job and fund of escrow job 3 (chore(contracts): deploy escrow on testnet and document verifiable on-chain evidence #84).
  • Ports: SubmissionLedger (Soroban adapter) and EscrowEffects (NoopEscrowEffects until feat(server): escrow relay prepare and submit endpoints for hires #96 implements it).
  • TrackerLoop: periodic, never overlapping, one session per pass, wired in server.dart behind PULS3_TRACKER_ENABLED (default false) and PULS3_TRACKER_INTERVAL_SECONDS (default 5). Documented in puls3_server/README.md.
  • Store follow-ups from the feat(server): chain submission store, sendTransaction and tracker for create_job and fund #100 review: JSON-RPC rejections typed separately from outages, unreadable rows isolated, recordSend only from submitted, purpose/preparation and error-code validation, and a single listSubmitted query ordered by a new lastCheckedAt so stuck records rotate instead of starving the rest.

Acceptance criteria (#97, demo scope)

  • A submission still pending on RPC (NOT_FOUND right after sending) stays submitted and is retried, never rejected
  • A server restart does not lose or duplicate tracking (integration test against Postgres)
  • A chain-confirmed funding whose effect or write fails is retried on the next pass, not lost
  • Reads never trigger a server-submitted call
  • release / claim_refund at their deadlines: after the demo (server-signed calls)

Verification evidence

Notes for reviewers

listSubmitted lists never-sent records first, then the least recently sent, so stuck records cannot starve the batch. An unreadable row is logged and failed with EscrowCallFailed instead of failing the whole batch. recordSend only counts sends of submitted records, insertSubmitted checks the purpose/preparation pairing and markFailed only accepts known outcome codes.
…escrow job 3

getTransaction with xdrFormat json from soroban-testnet.stellar.org for create_job 3baba183... and fund 43cd3e84... (contracts/deployments/testnet.json).
Each pass polls getTransaction for a batch of submitted records, resends the persisted envelope when due, expires records past their time bounds by chain time, and confirms createJob and fund through EscrowEffects. A record that cannot be read or updated stays submitted for the next pass without stopping the batch.
…nnot starve the tracker

listSubmitted put never-sent rows first and skipped the sent-rows query when they filled the batch, so limit or more records the tracker never sends (untracked purposes, SUCCESS without an effect) starved createJob and fund records. Its two queries could also list one row twice.

Add a server-only lastCheckedAt column (set on insert, backfilled from createdAt) and index (state, lastCheckedAt, id). listSubmitted is now one query ordered by lastCheckedAt then id, and the tracker calls the new conditional recordCheck for every record a pass leaves submitted. recordSend keeps pacing resends only.
@TOMOKI977 TOMOKI977 added this to the Stellar Elite (Oct 10) milestone Oct 5, 2026
@TOMOKI977 TOMOKI977 added area: backend Serverpod endpoints and persistence type: feat New functionality labels Oct 5, 2026
@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Deploying puls3 with  Cloudflare Pages  Cloudflare Pages

Latest commit: 86b2022
Status: ✅  Deploy successful!
Preview URL: https://cf3c58f2.puls3-4lw.pages.dev
Branch Preview URL: https://feat-97-tracker-loop.puls3-4lw.pages.dev

View logs

@moises-cisneros moises-cisneros left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Checked out the branch and ran it locally: dart analyze is clean and dart test passes (283 tests), including the integration suite and the concurrent confirm/fail test. I also started the real server with PULS3_TRACKER_ENABLED=true PULS3_TRACKER_INTERVAL_SECONDS=2 and the tracker starts (Chain submission tracker started, every 2s).

The design is good: conditional transitions, recordCheck rotation against starvation, skip-if-in-flight, and the flag off by default. A few things I would look at before merging:

  • Timeout on store calls (tracker_loop.dart). The RPC has an 8s timeout, but listSubmitted / recordCheck and session creation do not. A hung DB call leaves _inFlight set and every later tick is skipped silently. A pass-level timeout would avoid that.
  • No shutdown path (server.dart, chain_tracker_wiring.dart). The TrackerLoop returned by startChainTracker is discarded, so stop() is never called and the http.Client is never closed.
  • Config parsed after pod.start(). I ran it with PULS3_TRACKER_INTERVAL_SECONDS=abc: the web server is already listening when the ArgumentError crashes the process. It fails loudly with a clear message, so it is minor, but parsing before pod.start() would fail earlier.
  • Design question, NOT_FOUND after validUntil (chain_submission_tracker.dart). NOT_FOUND also covers hashes older than the RPC retention window. If the tracker is off for longer than that window, a tx that did land would be failed as PreparationExpired and the record cannot recover. I did not test this against a real node, so I am asking rather than claiming: is that window considered, or is a distinct code for "too old to know" worth it?
  • Design question, SUCCESS without a parsable event. It goes straight to final JobEvidenceUnavailable (covered by tests, so clearly intentional). Would a bounded retry be safer if a node answers SUCCESS before events are available?
  • If recordSend throws after a successful resend, the envelope is re-sent every pass. Harmless (DUPLICATE) but noisy.

Since NoopEscrowEffects marks confirmed with no domain effect, please keep the flag off until #96 lands, and add an idempotency test for the effects there.

…' into feat/97-tracker-loop

# Conflicts:
#	puls3_server/lib/src/chain/chain_submission_store.dart
#	puls3_server/test/support/chain_submission_store_contract.dart
A pass that never completed, such as one stuck on a database call, kept the loop's in-flight marker set, so every later tick was skipped. TrackerLoop now takes a pass timeout, logs a timed-out pass as an error and lets the next tick start a new pass. The wiring uses a five minute timeout.
server.dart dropped the loop startChainTracker returned, so it was never stopped and its HTTP client never closed. TrackerLoop now takes an onStop callback that runs after the pass in flight, the wiring uses it to close the client, and server.dart registers the loop's stop as a Serverpod shutdown task.
@TOMOKI977

Copy link
Copy Markdown
Contributor Author

Thanks for the thorough pass, especially starting the real server with the flag on.

Fixed on this branch (after merging the #100 fix in 89aae26):

  • Pass timeout (ed14fe8): TrackerLoop now takes a passTimeout (5 min in the wiring, since a full batch of up to 100 records with an 8s RPC limit each can legitimately take over a minute). A hung pass is logged as an error and the next tick runs. The hung future is abandoned, not cancelled; that's safe because every store transition only touches a record that is still submitted.
  • Shutdown path (86b2022): server.dart keeps the loop and registers tracker.stop with pod.experimental.shutdownTasks, so on SIGINT/SIGTERM Serverpod stops taking requests, waits for the pass in flight, then closes the http.Client via a new onStop hook before closing the DB. Note the API is marked experimental in Serverpod 3.4.13.

Suite: 285 passed (integration included).

Moving the rest to #101 so this PR stays focused:

  • Config parsed after pod.start(): agreed, parse before start.
  • NOT_FOUND after validUntil: you're right to ask. NOT_FOUND can't tell "never landed" from "older than the RPC retention window", so a long tracker outage could fail a tx that actually landed. The fix I have in mind is to stop treating that case as PreparationExpired when we're past the retention window, and reconcile from the escrow job state instead (or use a distinct "outcome unknown" code).
  • SUCCESS without a parsable event: a bounded retry before JobEvidenceUnavailable is safer, will do.
  • recordSend throwing after a resend: noisy but harmless, will look at it with the above.

And yes, PULS3_TRACKER_ENABLED stays off until #96 lands with real, idempotent effects.

@moises-cisneros moises-cisneros left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-checked the updated branch locally: dart analyze is clean and dart test passes (285 tests, integration included). I also started the real server with the flag on and the tracker starts as before.

  • The pass timeout (ed14fe8) and the shutdown path via shutdownTasks (86b2022) both look right. Abandoning a hung pass instead of cancelling it is fine given every transition only touches a submitted record, and the concurrent confirm/fail test backs that up.
  • The open design points (config parsed before pod.start(), NOT_FOUND past the retention window, bounded retry before JobEvidenceUnavailable) are tracked for #101, so I am not blocking on them. I will look at those there, especially how outcome-unknown is reconciled.

Approving. Keeping PULS3_TRACKER_ENABLED off until #96 lands, as you said.

@moises-cisneros
moises-cisneros merged commit ee3babd into feat/97-chain-submission-tracker Oct 6, 2026
1 check passed
TOMOKI977 added a commit that referenced this pull request Oct 6, 2026
… create_job and fund (#100)

* feat(server): add chain submission purpose, state and outcome code values

* feat(server): add sendTransaction to the Soroban RPC client

* feat(server): add chain_submission model and migration

* feat(server): add chain submission store contract and in-memory fake

* feat(server): add Serverpod chain submission store with conflict mapping

* fix(server): skip recordSend for submissions that are no longer submitted

recordSend locked the row but only checked that it exists, so a send recorded after a confirm or fail still bumped sendAttempts and lastSentAt of a final record. It now returns false unless the record is still submitted, in the Serverpod store and the in-memory fake alike.

* feat(server): chain submission tracker for create_job and fund (#102)

* fix(server): keep JSON-RPC error code and message and type request rejections

* fix(server): rotate, isolate and validate chain submission store records

listSubmitted lists never-sent records first, then the least recently sent, so stuck records cannot starve the batch. An unreadable row is logged and failed with EscrowCallFailed instead of failing the whole batch. recordSend only counts sends of submitted records, insertSubmitted checks the purpose/preparation pairing and markFailed only accepts known outcome codes.

* test(server): record the testnet create_job and fund transactions of escrow job 3

getTransaction with xdrFormat json from soroban-testnet.stellar.org for create_job 3baba183... and fund 43cd3e84... (contracts/deployments/testnet.json).

* feat(server): parse escrow job_created and job_funded events from transactions

* feat(server): add the submission ledger port and its Soroban RPC adapter

* feat(server): add the escrow effects port with a no-op implementation

* feat(server): add the chain submission tracker pass

Each pass polls getTransaction for a batch of submitted records, resends the persisted envelope when due, expires records past their time bounds by chain time, and confirms createJob and fund through EscrowEffects. A record that cannot be read or updated stays submitted for the next pass without stopping the batch.

* test(server): cover tracker recovery of persisted submissions after a restart

* feat(server): add a periodic tracker loop without overlapping passes

* feat(server): start the chain submission tracker when PULS3_TRACKER_ENABLED is true

* fix(server): list chain submissions by last check so stuck records cannot starve the tracker

listSubmitted put never-sent rows first and skipped the sent-rows query when they filled the batch, so limit or more records the tracker never sends (untracked purposes, SUCCESS without an effect) starved createJob and fund records. Its two queries could also list one row twice.

Add a server-only lastCheckedAt column (set on insert, backfilled from createdAt) and index (state, lastCheckedAt, id). listSubmitted is now one query ordered by lastCheckedAt then id, and the tracker calls the new conditional recordCheck for every record a pass leaves submitted. recordSend keeps pacing resends only.

* fix(server): time out a hung tracker pass so later passes run

A pass that never completed, such as one stuck on a database call, kept the loop's in-flight marker set, so every later tick was skipped. TrackerLoop now takes a pass timeout, logs a timed-out pass as an error and lets the next tick start a new pass. The wiring uses a five minute timeout.

* fix(server): stop the tracker and close its http client on shutdown

server.dart dropped the loop startChainTracker returned, so it was never stopped and its HTTP client never closed. TrackerLoop now takes an onStop callback that runs after the pass in flight, the wiring uses it to close the client, and server.dart registers the loop's stop as a Serverpod shutdown task.

* ci: allow Soroban ledger key hashes in gitleaks

Recorded getTransaction fixtures include key_hash, the public SHA-256 of a ledger key in a TTL entry, which the generic-api-key rule flags. The allowlist matches only that JSON field with a 64-character hex value.
@TOMOKI977
TOMOKI977 deleted the feat/97-tracker-loop branch October 6, 2026 18:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: backend Serverpod endpoints and persistence type: feat New functionality

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants