Repository navigation
feat(server): chain submission tracker for create_job and fund - #102
Conversation
listSubmitted lists never-sent records first, then the least recently sent, so stuck records cannot starve the batch. An unreadable row is logged and failed with EscrowCallFailed instead of failing the whole batch. recordSend only counts sends of submitted records, insertSubmitted checks the purpose/preparation pairing and markFailed only accepts known outcome codes.
…escrow job 3 getTransaction with xdrFormat json from soroban-testnet.stellar.org for create_job 3baba183... and fund 43cd3e84... (contracts/deployments/testnet.json).
Each pass polls getTransaction for a batch of submitted records, resends the persisted envelope when due, expires records past their time bounds by chain time, and confirms createJob and fund through EscrowEffects. A record that cannot be read or updated stays submitted for the next pass without stopping the batch.
…nnot starve the tracker listSubmitted put never-sent rows first and skipped the sent-rows query when they filled the batch, so limit or more records the tracker never sends (untracked purposes, SUCCESS without an effect) starved createJob and fund records. Its two queries could also list one row twice. Add a server-only lastCheckedAt column (set on insert, backfilled from createdAt) and index (state, lastCheckedAt, id). listSubmitted is now one query ordered by lastCheckedAt then id, and the tracker calls the new conditional recordCheck for every record a pass leaves submitted. recordSend keeps pacing resends only.
Deploying puls3 with
|
| Latest commit: |
86b2022
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://cf3c58f2.puls3-4lw.pages.dev |
| Branch Preview URL: | https://feat-97-tracker-loop.puls3-4lw.pages.dev |
moises-cisneros
left a comment
There was a problem hiding this comment.
Checked out the branch and ran it locally: dart analyze is clean and dart test passes (283 tests), including the integration suite and the concurrent confirm/fail test. I also started the real server with PULS3_TRACKER_ENABLED=true PULS3_TRACKER_INTERVAL_SECONDS=2 and the tracker starts (Chain submission tracker started, every 2s).
The design is good: conditional transitions, recordCheck rotation against starvation, skip-if-in-flight, and the flag off by default. A few things I would look at before merging:
- Timeout on store calls (
tracker_loop.dart). The RPC has an 8s timeout, butlistSubmitted/recordCheckand session creation do not. A hung DB call leaves_inFlightset and every later tick is skipped silently. A pass-level timeout would avoid that. - No shutdown path (
server.dart,chain_tracker_wiring.dart). TheTrackerLoopreturned bystartChainTrackeris discarded, sostop()is never called and thehttp.Clientis never closed. - Config parsed after
pod.start(). I ran it withPULS3_TRACKER_INTERVAL_SECONDS=abc: the web server is already listening when theArgumentErrorcrashes the process. It fails loudly with a clear message, so it is minor, but parsing beforepod.start()would fail earlier. - Design question,
NOT_FOUNDaftervalidUntil(chain_submission_tracker.dart).NOT_FOUNDalso covers hashes older than the RPC retention window. If the tracker is off for longer than that window, a tx that did land would be failed asPreparationExpiredand the record cannot recover. I did not test this against a real node, so I am asking rather than claiming: is that window considered, or is a distinct code for "too old to know" worth it? - Design question,
SUCCESSwithout a parsable event. It goes straight to finalJobEvidenceUnavailable(covered by tests, so clearly intentional). Would a bounded retry be safer if a node answersSUCCESSbefore events are available? - If
recordSendthrows after a successful resend, the envelope is re-sent every pass. Harmless (DUPLICATE) but noisy.
Since NoopEscrowEffects marks confirmed with no domain effect, please keep the flag off until #96 lands, and add an idempotency test for the effects there.
…' into feat/97-tracker-loop # Conflicts: # puls3_server/lib/src/chain/chain_submission_store.dart # puls3_server/test/support/chain_submission_store_contract.dart
A pass that never completed, such as one stuck on a database call, kept the loop's in-flight marker set, so every later tick was skipped. TrackerLoop now takes a pass timeout, logs a timed-out pass as an error and lets the next tick start a new pass. The wiring uses a five minute timeout.
server.dart dropped the loop startChainTracker returned, so it was never stopped and its HTTP client never closed. TrackerLoop now takes an onStop callback that runs after the pass in flight, the wiring uses it to close the client, and server.dart registers the loop's stop as a Serverpod shutdown task.
|
Thanks for the thorough pass, especially starting the real server with the flag on. Fixed on this branch (after merging the #100 fix in 89aae26):
Suite: 285 passed (integration included). Moving the rest to #101 so this PR stays focused:
And yes, |
moises-cisneros
left a comment
There was a problem hiding this comment.
Re-checked the updated branch locally: dart analyze is clean and dart test passes (285 tests, integration included). I also started the real server with the flag on and the tracker starts as before.
- The pass timeout (ed14fe8) and the shutdown path via
shutdownTasks(86b2022) both look right. Abandoning a hung pass instead of cancelling it is fine given every transition only touches asubmittedrecord, and the concurrent confirm/fail test backs that up. - The open design points (config parsed before
pod.start(),NOT_FOUNDpast the retention window, bounded retry beforeJobEvidenceUnavailable) are tracked for #101, so I am not blocking on them. I will look at those there, especially how outcome-unknown is reconciled.
Approving. Keeping PULS3_TRACKER_ENABLED off until #96 lands, as you said.
ee3babd
into
feat/97-chain-submission-tracker
… create_job and fund (#100) * feat(server): add chain submission purpose, state and outcome code values * feat(server): add sendTransaction to the Soroban RPC client * feat(server): add chain_submission model and migration * feat(server): add chain submission store contract and in-memory fake * feat(server): add Serverpod chain submission store with conflict mapping * fix(server): skip recordSend for submissions that are no longer submitted recordSend locked the row but only checked that it exists, so a send recorded after a confirm or fail still bumped sendAttempts and lastSentAt of a final record. It now returns false unless the record is still submitted, in the Serverpod store and the in-memory fake alike. * feat(server): chain submission tracker for create_job and fund (#102) * fix(server): keep JSON-RPC error code and message and type request rejections * fix(server): rotate, isolate and validate chain submission store records listSubmitted lists never-sent records first, then the least recently sent, so stuck records cannot starve the batch. An unreadable row is logged and failed with EscrowCallFailed instead of failing the whole batch. recordSend only counts sends of submitted records, insertSubmitted checks the purpose/preparation pairing and markFailed only accepts known outcome codes. * test(server): record the testnet create_job and fund transactions of escrow job 3 getTransaction with xdrFormat json from soroban-testnet.stellar.org for create_job 3baba183... and fund 43cd3e84... (contracts/deployments/testnet.json). * feat(server): parse escrow job_created and job_funded events from transactions * feat(server): add the submission ledger port and its Soroban RPC adapter * feat(server): add the escrow effects port with a no-op implementation * feat(server): add the chain submission tracker pass Each pass polls getTransaction for a batch of submitted records, resends the persisted envelope when due, expires records past their time bounds by chain time, and confirms createJob and fund through EscrowEffects. A record that cannot be read or updated stays submitted for the next pass without stopping the batch. * test(server): cover tracker recovery of persisted submissions after a restart * feat(server): add a periodic tracker loop without overlapping passes * feat(server): start the chain submission tracker when PULS3_TRACKER_ENABLED is true * fix(server): list chain submissions by last check so stuck records cannot starve the tracker listSubmitted put never-sent rows first and skipped the sent-rows query when they filled the batch, so limit or more records the tracker never sends (untracked purposes, SUCCESS without an effect) starved createJob and fund records. Its two queries could also list one row twice. Add a server-only lastCheckedAt column (set on insert, backfilled from createdAt) and index (state, lastCheckedAt, id). listSubmitted is now one query ordered by lastCheckedAt then id, and the tracker calls the new conditional recordCheck for every record a pass leaves submitted. recordSend keeps pacing resends only. * fix(server): time out a hung tracker pass so later passes run A pass that never completed, such as one stuck on a database call, kept the loop's in-flight marker set, so every later tick was skipped. TrackerLoop now takes a pass timeout, logs a timed-out pass as an error and lets the next tick start a new pass. The wiring uses a five minute timeout. * fix(server): stop the tracker and close its http client on shutdown server.dart dropped the loop startChainTracker returned, so it was never stopped and its HTTP client never closed. TrackerLoop now takes an onStop callback that runs after the pass in flight, the wiring uses it to close the client, and server.dart registers the loop's stop as a Serverpod shutdown task. * ci: allow Soroban ledger key hashes in gitleaks Recorded getTransaction fixtures include key_hash, the public SHA-256 of a ledger key in a TTL entry, which the generic-api-key rule flags. The allowlist matches only that JSON field with a 64-character hex value.
Closes #97 (scope for the Stellar Elite demo:
create_jobandfundconfirmation)Summary
The durable tracker from the relay design (
docs/architecture/api.md, steps 4 and 5):ChainSubmissionTracker.pass(), persubmittedrecord:NOT_FOUND: if the chain's close time is pastvalidUntil, it fails withPreparationExpired; otherwise it resends the persisted envelope everyresendAfter(30 s). A rejected envelope fails withSubmissionRejected.FAILED: fails withTransactionFailed.SUCCESSforcreateJob/fund: needs the matching escrow event (elseJobEvidenceUnavailable), then callsEscrowEffects.EffectOkconfirms the record; a failed effect fails it withJobMismatch/JobEvidenceUnavailable.LedgerUnavailableor an exception leaves the recordsubmittedfor the next pass.job_createdandjob_fundedparsed fromgetTransaction, filtered by the configured escrow contract id. Fixtures are the real testnetcreate_jobandfundof escrow job 3 (chore(contracts): deploy escrow on testnet and document verifiable on-chain evidence #84).SubmissionLedger(Soroban adapter) andEscrowEffects(NoopEscrowEffectsuntil feat(server): escrow relay prepare and submit endpoints for hires #96 implements it).TrackerLoop: periodic, never overlapping, one session per pass, wired inserver.dartbehindPULS3_TRACKER_ENABLED(defaultfalse) andPULS3_TRACKER_INTERVAL_SECONDS(default 5). Documented inpuls3_server/README.md.recordSendonly fromsubmitted, purpose/preparation and error-code validation, and a singlelistSubmittedquery ordered by a newlastCheckedAtso stuck records rotate instead of starving the rest.Acceptance criteria (#97, demo scope)
NOT_FOUNDright after sending) stayssubmittedand is retried, never rejectedrelease/claim_refundat their deadlines: after the demo (server-signed calls)Verification evidence
dart analyze --fatal-infos: clean (server and client).dart test: 254 unit and 29 integration tests pass (integration against Docker Postgres).serverpod generateleaves no diff. One new migration on top of feat(server): chain submission store, sendTransaction and tracker for create_job and fund #100's (20261005231322835, additive plus a backfill).createJob/fundinlistSubmitted), corroborated, fixed in90da30band approved by a scoped validator. Non-blocking findings are tracked in chore(server): harden the chain submission tracker after the Stellar Elite demo #101.Notes for reviewers
EscrowEffects(lib/src/chain/escrow_effects.dart) is the hook for feat(server): escrow relay prepare and submit endpoints for hires #96. It must be idempotent, because a crash between the effect andmarkConfirmedruns it again, and it should verify that the event's job matches the hire (only the first matching event is passed).fundthat succeeded on chain but failed verification record "funds may have moved" on the hire (PaymentAlreadySubmitted)?