Skip to content

Echo - #25

Merged
dundich merged 67 commits into
mainfrom
echo
Sep 29, 2026
Merged

Echo#25
dundich merged 67 commits into
mainfrom
echo

Conversation

@dundich

@dundich dundich commented Sep 29, 2026

Copy link
Copy Markdown
Owner

No description provided.

dundich and others added 30 commits September 14, 2026 23:02
… check exit code in GetChannelsAndSampleRate
…view fixes

Signed-off-by: dundich <dundich1@gmail.com>
- Single-flight dispose gate (RunDispose/RunDisposeAsync): the winner runs
  shutdown and disposes _shutdownCts, concurrent losers wait for it
- Bounded reader wait on shutdown (sync + async) via ShutdownTimeout;
  reader loop exits quietly on ObjectDisposedException instead of a
  spurious ShutdownAsync and 'ReaderTask error' log
- Merge byte-identical ClearRemainingItems into DrainAndResetIdle
- RandomIndices: rejection sampling of distinct indices instead of the
  zero-filled partial Fisher-Yates that always picked reader #0
- Drop the _readerCount counter in favour of _ctsReaders.Count; extract
  shared CancelAndTrackReaders
- Route the duplicated reader-timeout log through ThrowHelper.ReadersTimeout
- Docs: DrainAndResetIdle naming and _ctsReaders.Count references
- Test: ConcurrencyLimit=0 pause - enqueue is accepted, waits, processes
  after the limit is restored
…Many

Introduce SaEnqueueStrategy (Wait/Skip/Throw) governing bounded-buffer
behavior, non-blocking TryEnqueue, batch EnqueueMany, Skipped status,
SaWorkQueueFullException and AvailableCapacity. Refactor logging into a
SaWorkQueueLogMessages source generator and extract ThrowHelper.

Update Readme.md / Readme-ru.md with the new features.
Correctness
- DrainAndResetIdle: report Aborted for channel items whose caller token
  was already cancelled (previously they vanished without any terminal
  status), and call MarkInactive on that path
- DrainAndResetIdle: zero the pending count only when every reader task
  has completed, so IsIdle()/QueueTasks stay honest when a reader
  survives the bounded shutdown wait
- _state: reads now go through Volatile.Read (writes already used
  Interlocked) for correct memory-model visibility

Refactor
- ExecuteItemAsync: skip the per-item linked CTS when the caller token
  cannot be cancelled (the common default/None path)
- RandomIndices: partial Fisher-Yates instead of rejection sampling
  (O(k), guaranteed-distinct, uniform)

Docs
- ISaWorkQueue: TimeoutException for ForceCancelReadersAsync timeout
  (was documented as OperationCanceledException), OperationCanceledException
  for Enqueue (Wait strategy), clamp note on ConcurrencyLimit, shutdown
  wording aligned with actual behaviour (cancels readers, does not
  finish in-flight work)
- SaWorkQueueOptions: WithConcurrencyLimit 0 = pause (was "unlimited"),
  HandleItemFaulted default = ShutdownQueue
- Setup.CreateSimple: int? concurrency = null instead of magic -1, + XML docs
- Readme.md / Readme-ru.md: queue capacity default is MaxConcurrency (not
  the limit), IsIdle() is a method (was property), ShutdownAsync interrupts
  in-flight work (was "finish active")

Tests
- New WorkQueueDrainTests: regression tests for the two drain fixes,
  verified to fail on the previous code
- WorkQueueConcurrencyTests: helper loop no longer swallows non-cancellation
  exceptions into the console
- .gitattributes: * text=auto eol=lf (*.bat/*.cmd keep CRLF)
- .editorconfig: end_of_line = lf, insert final newline
- .vscode/settings.json: files.eol = \n
- .gitignore: drop root .vscode/ rule so the !.vscode/*.json negations actually work
Signed-off-by: dundich <dundich1@gmail.com>
Sa.Configuration:
- Arguments: params string[] ctor, negative numbers parsed as values (IsFlag), drop IsPresent
- ChainedSecretStore: thread-safe LIFO chain (Interlocked.Exchange, last store wins)
- SecretService: depth-based circular-reference guard (max depth 3)
- ISecretService: add GetSecret; Secrets.CreateDefault normalizes FileName/Args/EnvironmentName
- SecretOptions: immutable record with nullable Args/EnvironmentName
- FileSecretStore: trim lines, skip lines without '='
- csproj: Microsoft.Extensions.Configuration + Hosting.Abstractions (was Microsoft.Extensions.Hosting)

Sa.Configuration.PostgreSql:
- provider: reload replaces the full key set (stale keys drop), swap under Lock
- csproj: direct Npgsql reference

Tests: update Arguments/Secrets/SecretService/PostgreSql suites, add FileSecretStoreTests
Docs: sync Readme/Readme-ru (EN/RU) with the new behavior
~
Signed-off-by: dundich <dundich1@gmail.com>
…ule unit tests

- src/Directory.Build.props: point RestoreFallbackFolders at a real
  directory on non-Windows so PublishAOT/analyzer projects don't fail
  with MSB4018 (bogus C:\Program Files (x86) fallback path) on WSL/Linux
- .gitignore: ignore src/.nuget and src/.localpilot
- WorkQueue.Console: add WSL2 launch profile
- JobScheduler.DisposeAsync: await ctsStopping.CancelAsync() so stopping
  jobs are released and their DI scopes don't leak
- Sa.ScheduleTests: add JobExecutorTests, JobControllerLifecycleTests and
  a TrackingScopeFactory helper (real DI container, no mocking framework)
- Sa.HybridFileStorageTests: use a synchronous IProgress<> in the
  CopyToScopeBatchAsync progress test to remove the SynchronizationContext
  timing dependency of the built-in Progress<T>
- Sa.Outbox.PostgreSqlTests: PauseAndResume test now relies on the
  ProcessMessages return value instead of a static counter shared across
  parallel tests publishing to the same container
…se/Resume test

- Skip 5 CrossFeedSeparatorIntegrationTests due to missing data/pcm_s16le.wav
- Fix Manager_PauseAndResume_MessagesProcessedAfterResume flaky test:
  - Add settling delay before pause to let in-flight polling finish
  - Re-read settings from manager after Pause/Resume to get latest snapshot
  - Verify Paused flag is reflected in retrieved settings
dundich added 28 commits September 27, 2026 23:04
…ns, idle-wait semantics, options validation

- B1: RandomIndices Fisher-Yates over a full-sized buffer; any limit
  decrease with toCancel < live used to throw IndexOutOfRangeException
  out of the ConcurrencyLimit setter
- B2/B5: force-cancel accounting is eager and idempotent inside the
  cancel critical section; RemoveReader only decrements untracked
  terminations, so dying readers no longer eat a limit re-armed by
  JobScheduler.AbortJob -> Start
- B3: DrainAndResetIdle takes a SaWorkDrainReason; dropped items report
  "force-cancelled" instead of "shut down" (status stays Faulted)
- B4: Shutdown/ShutdownAsync split into per-step try/catch with
  Writer.TryComplete and drain in finally — a faulting cancellation
  callback no longer leaves the channel open and IsIdle() false forever
- B6: documented that Shutdown/ForceCancel/Dispose must not be called
  from ISaWork.Execute (blocks for the full ShutdownTimeout)
- B7: explicit _paused flag + WaitForIdleAsync(failIfPaused) instead of
  the _concurrency == 0 heuristic; new guard for an active queue with no
  live readers (returns/throws instead of hanging on unreachable idle)
- B8: SaWorkQueue ctor re-validates ConcurrencyLimit >= 0, QueueCapacity
  >= 1, positive ShutdownTimeout (record ctor bypasses With* guards)
- D1-D7: docs corrections, [DoesNotReturn], TryAdd in DI setup, init
  exception props, null checks, dead NoWarn removed
- tests: cancellation-order theories, force-cancel/limit re-arm race,
  failIfPaused and no-readers wait semantics, drain-reason text,
  faulting-callback shutdown, options validation
- docs: Readme/Readme-ru synced; sa-utils-workqueue-improvements.md

Version stays 0.12.0.
… all

- C1 (critical): StartReaderUnderLock called ReaderLoopAsync BEFORE
  adding the CTS to _ctsReaders/_ctsWorks/_taskReaders. With a
  non-empty buffer and a processor that completes or faults without
  yielding, the loop ran to completion inline and hit
  RemoveReader(cts) with IndexOf == -1. System.Threading.Lock is
  re-entrant, so the nested call was not blocked and silently
  mis-accounted an unregistered reader: the requested limit was
  eaten (asked 4, got 0), a permanently disposed CTS was appended
  (it reports IsCancellationRequested == false, so HasLiveReaders()
  lied), CancelReadersUnderLock then called Cancel() on it and threw
  ObjectDisposedException out of the public ConcurrencyLimit setter,
  WaitForIdleAsync hung forever on an unreachable idle (failIfPaused
  did not help, the heuristic still saw readers), and the work CTS was
  never disposed. Reachable on the DEFAULT path: ShutdownQueue or
  StopReader + a paused queue with a full buffer + ConcurrencyLimit = N
- fix: StartReaderUnderLock returns early when task.IsCompleted, an
  exact signal (a task cannot complete before its finally ran, and a
  completion on another thread blocks on _readersSync, still held)
- fix: RemoveReader(cts, ctsWork) no longer touches _concurrency,
  _pendingRemovals, _intentionalRemovals, _forceCancelled or the
  ReaderLost log when the reader was never registered; taking ctsWork
  as a parameter lets it dispose the work CTS that nothing else sees
- _concurrency now reflects the requested limit, HasLiveReaders() the
  real one — consistent with the StopReader contract and with the
  same argument that fixed B2: a dying reader must not eat a limit
  the caller set
- D1: the "Called without lock" comment in OnStatusChanged was false.
  A reader started over a non-empty buffer reports synchronously
  inside StartReaderUnderLock, so the handler runs under
  _readersSync; documented the deadlock constraint, plus the
  re-entrancy of Lock on the lock fields
- tests: WorkQueueSyncReaderTests (9) — limit survives a synchronous
  fault on both StopReader and the default ShutdownQueue path, no dead
  reader left behind, no ObjectDisposedException out of the setter,
  WaitForIdle reaps instead of hanging, failIfPaused sees the empty
  pool and recovers, 5 re-arm rounds accumulate nothing, plus the
  healthy suspending-processor and 2->5->1->4 scaling regressions
- docs: sa-utils-workqueue-improvements.md rewritten for the current
  state — C1 with reproduction and fix, open M1-M4/L1-L8/D2-D3, a
  "checked and NOT a bug" section, .NET 10 notes. M2 corrected: a
  parked producer is released by Shutdown (via Writer.TryComplete)
  and by raising ConcurrencyLimit, not only by the caller token, so
  linking the wait to _shutdownCts would be redundant

Solution build 0 warnings/0 errors; Sa.Utils.WorkQueue.Tests 108
(was 99, 5 runs without flakiness); Sa.ScheduleTests 133.

Version stays 0.12.0.
…arantees

Phase 0 of sa-utils-workqueue-refactor-plan.md: turn the seven
"Critical Rules for AI Agents" out of the readme prose and into
structure, so the rules stop needing to be remembered. The rules
were never about discipline, they were descriptions of an
implementation nobody could see from the outside.

- split SaWorkQueue.cs (1129 lines) into four partials by concern:
  .Readers (pool), .Shutdown (force-cancel, drain, dispose),
  .Enqueue (producers), and the remainder (state, dispatch, public
  surface)
- replace _ctsReaders/_ctsWorks/_taskReaders/_intentionalRemovals/
  _forceCancelled/_pendingRemovals with one sealed Reader record
  carrying a Task and a ReaderStop reason (None/LimitDecrease/
  ForceCancel). Six parallel lists had to agree; the type makes
  disagreement unrepresentable
- StartReaderUnderLock now registers the reader BEFORE starting its
  loop, so a reader that dies during startup is indistinguishable
  from one that dies later. 32 lines to 11, no branches. This
  supersedes the task.IsCompleted guard added in c83c524 and
  removes the idx < 0 branch from RemoveReader (74 lines to 28)
  instead of documenting it. Behaviour change: re-arming over a
  non-empty buffer with a synchronously failing processor now
  honestly drops the limit (asked 4, got 0) instead of keeping 4.
  Sa.Schedule is unaffected — it re-sets the limit on every
  AbortJob/Start
- L5: _taskCount is guarded by _pendingSync and read through
  Volatile.Read; it is no longer both volatile and locked
- L6: the disposed/stopped prologue is one method,
  ThrowIfNotActive(), instead of four copies
- locks are now named for what they protect (_readersSync for the
  pool, _pendingSync for the pending count and idle signal), and
  both fields carry the re-entrancy note that explains why the
  inline RemoveReader is harmless

Phase 1, steps 1/2/4 — all three about a diagnosis reaching its
addressee:

- M4: ExecuteItemAsync caught OperationCanceledException ex and
  then called OnStatusChanged(item, Aborted) / (..., Cancelled)
  WITHOUT it, unlike Faulted which always carries one. A status
  callback has no logger, so ex was the only account it could get
  of why an item stopped. Now passes the caught ex — not a
  synthetic ThrowHelper.CallerCancelledException(), which belongs
  in DrainAndResetIdle where no real exception exists
- L3: ThrowHelper.QueueStopped(Exception? cause = null) forwards
  the channel exception that woke a parked producer, so callers
  stop seeing a bare "ChannelClosedException"
- L4: HandleShutdownOnError wrote _shutdownError unconditionally.
  ShutdownAsync returns immediately once the state is taken, so
  every later fault raced to overwrite the root cause. Now ??=,
  and the field has a single write site
- L1: ConcurrencyLimit = -1 silently paused the queue forever while
  IsEnabled stayed true — a failure with no symptom. The
  constructor and WithConcurrencyLimit already rejected it; the
  setter was the only hole. Now throws. Clamping is kept only for
  a positive value above MaxConcurrency, where the bound is
  meaningful and the value must not be discarded
- L7: AddSaWorkQueue accepted Scoped/Transient, which the readme
  forbade in prose only. Each resolution then got its own reader
  pool and its own copy of the buffer, working right up until two
  scopes enqueue the same work. Now rejected at registration;
  the <TProcessor, TInput> overload was already hard-wired
  Singleton, so both paths now agree
- L8: MaxConcurrency < 1 was folded to ProcessorCount in
  silence. 0 and null stay "processor count" (documented by
  WithMaxConcurrency(0)); a negative value has no reading and is
  rejected by the constructor next to the other validations
- D3: failIfPaused -> failIfNoProgress. The flag threw both
  QueuePaused and QueueHasNoReaders, which differ only in message
  text, so splitting them by flag would just make callers learn
  internals. The XML doc also asserted the opposite of the code
  ("reader loss is not a pause, the wait keeps waiting") for a
  branch that returns; rewritten to the actual behaviour
- D2: removed with the AI-rules section; the StopReader row in the
  error-strategy table now says what the code does

- tests: WorkQueueSyncReaderTests (10) rewritten for the new
  structure and pinned to the private _readers list and the Reader
  Loop/Work/IsLive properties, plus
  ReArm_WhileCancellingReaders_DoesNotLetThemEatTheNewLimit which
  holds old rule 5 directly. WorkQueueDiagnosticsTests (7) covers
  M4/L3/L4/L1/L7/L8. Two existing tests were rewritten to assert
  the new contract instead of the clamp
- every wave-2 fix was verified by reverting it: with the old code
  restored exactly four tests fail — both M4 statuses, L3 and L4.
  The first L4 attempt passed on the old code and was therefore
  proving nothing; it is now deterministic (two items in flight,
  the second ignoring cancellation and faulting only after the
  test has seen ShutdownError written)
- docs: sa-utils-workqueue-improvements.md restated for the
  current state with closed findings moved under "what was done"
  and a rationale for each; new
  sa-utils-workqueue-refactor-plan.md carries the phase-0
  rule-to-mechanism table and the remaining steps. Both readmes
  updated for the DI lifetime, the validation rules and the
  renamed parameter

Still open, and why: M3 (ForceCancelReaders* has no try/finally,
so a timeout wedges _taskCount and IsIdle() contradicts
WaitForIdleAsync) is waiting on a product decision about drain
semantics; M1 (the status callback runs under _readersSync and
deadlocks) needs the lock scope restructured, so it goes last.

Solution build 0 warnings/0 errors; Sa.Utils.WorkQueue.Tests 117
(was 108, 5 runs without flakiness); Sa.ScheduleTests 133.

Version stays 0.12.0. D3 is the only caller-visible break, and the
package is pre-1.0 — versions here are bumped repo-wide in a single
ticket, not per fix.
…y tested, then close M3

Two things, in the order they had to happen: the infrastructure that
made M3 reproducible, and M3 itself.

TimeProvider injection

ShutdownTimeout defaults to 30s, and every path that waits on it — a
force-cancel, a shutdown, a processor that ignores cancellation — was
reachable only by sleeping through that. So the M3 test could not be
written at all, let alone written deterministically.

- SaWorkQueueOptions.TimeProvider + WithTimeProvider, defaulting to
  TimeProvider.System. It governs waits only, never the processors, so
  a clock that jumps cannot corrupt a queue's state
- all four bounded waits moved onto it. Task.WaitAll(tasks, timeout)
  has no TimeProvider overload, so the two sync paths are a timed
  WaitAsync blocked on rather than awaited: blocking on the task, not
  via await, keeps the caller's synchronization context out of a
  thread that is already parked
- ManualTimeProvider in the test project, hand-written rather than
  pulled from Microsoft.Extensions.TimeProvider.Testing — a 30-second
  dependency for four timer-backed waits is a poor trade. Its
  ArmedCount lets a test wait until the code under test has really
  parked on a wait instead of guessing with a sleep; without it the
  advance can land before the timer exists and be lost
- WorkQueueTimeoutTests (6) assert the wait sits on the injected
  clock: parked while frozen, released by Advance. Reverting the
  injection fails four of them

M3 — the drain skipped when the readers do not unwind

ForceCancelReadersAsync had no try/finally, so an elapsed timeout or a
cancelled ct escaped before DrainAndResetIdle: the buffer and the
pending count survived, IsIdle() answered false for ever, and
WaitForIdleAsync returned as if all was well. Two APIs, one
situation, opposite answers.

The audit had this wider than it was. Task.WaitAll(tasks, timeout)
returns bool and does not throw TimeoutException, and a reader task
cannot fault — ReaderLoopAsync catches everything — so the
synchronous ForceCancelReaders was never affected, nor was
Shutdown. One method, and the cancelled-ct trigger was missing from
the finding entirely, which is the more dangerous one: a caller with
a short-lived token lost the drain silently.

Resolved as: keep the buffer. Nobody rejected those items and a pool
re-armed afterwards can still take them, so discarding work nobody
asked to discard is the worse failure. The synchronous variant is
untouched — there, "stop the pool and throw away the rest" is what
the operation means, and an explicit timeout is a statement about
how long readers should take, not permission to drop the queue.

- ForceCancelReadersAsync catches the timeout and the cancellation,
  logs the shortfall with the counts that matter, and rethrows so
  the caller learns the wait expired instead of inferring success
- WaitForIdleAsync logs a warning on both no-progress returns, naming
  the cause (Paused or NoReaders). IsIdle() and WaitForIdleAsync
  stop contradicting each other: the first answers "is there work",
  the second "is waiting still worth it", and the second now says so
  out loud
- contract written out in ISaWorkQueue.ForceCancelReadersAsync

WorkQueueForceCancelDrainTests (5). The end-to-end one carries it to
the finish: expired wait, the buffered item never marked Faulted,
both APIs admitting the state, a re-armed pool processing what the
abandoned one did not. Reverting the fix fails three; the other two
are deliberate boundary guards (an unbounded wait still drains, and a
queue that is draining stays quiet) and pass either way.

Incidental finding, not a bug: an item whose reader was cancelled but
whose processor returned on its own completes as Completed, not
Cancelled. The work really did run. Worth knowing when reading
statuses.

Sa.Utils.WorkQueue.Tests 129/129 (5 runs), Sa.ScheduleTests 133/133,
solution builds with 0 warnings.
… the status callback deadlocks

The last open finding. A reader started over a non-empty buffer picks
its item up on the spot and reports it through the status callback —
so StartReadersUnderLock ran caller-supplied code while holding
_readersSync. A handler that hands its event to another thread and
waits for it stood no chance: that thread needed the very lock the
handler's own thread was holding. Reproduced, and it was documented in
OnStatusChanged rather than fixed, because the fix was a restructuring
of the locking and deserved its own change.

Registration and launch are now separate steps:

- RegisterReaderUnderLock creates a reader, puts it in the registry and
  publishes Reader.Completion. No loop exists yet, so the loop's task
  cannot be published — but the completion source is, and LaunchReader
  completes it when the loop ends
- LaunchReader runs the loop and bridges its completion. The inline
  case — the loop ends before it returns — completes synchronously, so
  start-then-immediately-finish stays indistinguishable from
  start-then-finish-a-millisecond-later
- StartReaders registers under the lock and launches after releasing
  it. The decisions stay under the lock, so LiveReaderCount is still
  read atomically with the limit change and two concurrent changes
  cannot both decide to create the same slot

Reader.Task is no longer nullable. It could not be otherwise once
registration and launch were split, and the nullable version had every
observer asking a question it should not have to: is this reader the
one without a task, or the one about to get one? SnapshotReaderTasks
loses its OfType<Task>(), IsFinished loses its null caveat, and the
window is now unrepresentable rather than guarded.

Inline execution is kept deliberately. A reader still takes its first
item on the thread that resized the pool, which is not an optimisation:
Sa.Schedule's pause gate is built on it (JobController.WaitIfPaused
yields rather than blocking precisely because the loop runs on the
caller's thread until the first real yield).

Re-entrancy of System.Threading.Lock is no longer load-bearing, and
that is worth stating rather than leaving the docs claiming it. C1
rested on it — a nested RemoveReader from an inline loop. That case is
now impossible in principle, since loops are launched outside the
lock. Register-before-launch is still load-bearing: the reader is in
the registry before it can leave it.

The residual constraint is documented, not hidden: the callback still
runs on the reader's thread, so a handler that waits for that same
reader is waiting on itself. That is not about locks, it is what inline
execution means. Both readmes said the callback runs "on a thread-pool
thread", which was never true; they now say which thread it is and
what the one forbidden thing is.

WorkQueueStatusCallbackTests. The first is the deadlock itself: paused
queue, buffered item, a processor that never yields, and a callback
that touches the pool from another thread and waits. Reverting the fix
— putting the launch back under the lock — fails exactly that one, in
10s, on the probe's deadline rather than hanging the run. The second
pins inline execution on the caller's thread and passes either way: it
guards a property the fix could have broken, not a defect it removed.

Also closed M2's documentation, the only part of it outstanding. The
readmes described the paused queue from WaitForIdleAsync's side and
said nothing about the producer parked by Enqueue(Wait) on a full
buffer: released by a positive ConcurrencyLimit, a shutdown, or the
caller's own token, deliberately, because parking bounds the producer
and refusing the item would turn a pause into data loss.

sa-utils-workqueue-refactor-plan.md is now closed — phase 0, all five
steps of phase 1, and the TimeProvider injection. L2 and P5 stay out of
the plan and need their own tickets.

Sa.Utils.WorkQueue.Tests 131/131 (5 runs), Sa.ScheduleTests 133/133,
solution builds with 0 warnings.
…gnored, not recorded

L2, and the last open finding. SetConcurrencyLimit wrote _concurrency
and _paused before the IsEnabled check, so after a shutdown a caller
could assign 4, read 4 back, and conclude that four readers were on
their way. None were. The getter reported a number the pool did not
have, and a caller that only writes and reads had no way to notice.

The assignment is a statement of intent, not an operation on a
resource the caller believes it holds, so ignoring it silently is
right and throwing is wrong: the common caller is a shared
configuration path that does not branch on the queue's state, and an
exception there would be a failure the caller cannot act on. What
makes that safe is that the getter keeps reporting a number that was
true — the limit the queue was stopped at.

Note that number is not always the limit it was stopped with:
DisposeAsync takes it to zero on its way through reader teardown.
That is what "the limit it was stopped at" has to mean, and the
contract now says so instead of leaving it to be discovered.

The second effect is the one I did not go looking for. _paused no
longer becomes true because someone assigned zero to a stopped
queue. It used to, and WaitForIdleAsync would then name the reason
Paused for a queue that is stopped rather than paused — and
failIfNoProgress would throw QueuePaused where QueueHasNoReaders is
the truth. Stopped and paused read alike through that flag, which is
the same class of lie this commit removes from the limit.

The race stays and is not a consequence of the decision: assign the
limit concurrently with a shutdown and the IsEnabled check can pass
first, and the assignment is lost regardless. What the no-op makes
explicit is only the already-stopped case.

Three tests in WorkQueueConcurrencyTests. The two that carry the fix
compare the getter before and after the assignment; reverting the
check's position fails exactly those two. The third asserts the limit
still applies on a live queue and that the pool follows — a queue
ignoring every assignment would satisfy "changes nothing after
shutdown" perfectly while being broken in the only case that matters.
That one passes either way, deliberately.

The audit claimed L2 had no tests and reported 117/117; both were
stale by the time this landed, and are corrected. P2 and P4 were
listed as open in the same document after phase 0 had closed them.

Contract written out in ISaWorkQueue.ConcurrencyLimit, including the
negative-value rejection that was previously only in the
implementation, plus both readmes.

Sa.Utils.WorkQueue.Tests 134/134 (8 runs across the change),
Sa.ScheduleTests 133/133, solution builds with 0 warnings.
…d one stop looking alike

P5, the last finding open, plus a signature fix that came with it.
Neither is separable in the tree: both touch ISaWorkQueue, both need
the whole test suite, and splitting them would mean committing a state
that does not compile.

P5 — the finding was right, the proposed API was not. The audit asked
for `int LiveReaders`. Force-cancel releases the slots of the readers
it stops, so after one the limit reads 0 — and the live count reads 0
too. A deliberate pause gives 0 and 0. All three cases P5 exists to
tell apart produce the same two numbers, which is why a count was no
better than nothing: it repeated a number the queue already published.

So: SaWorkPoolState PoolState — Active, Paused, NoReaders, Stopped.
One named answer instead of a pair of fields the caller has to combine
correctly, and something that can go into a log line or a metric.

Stopped was added beyond the three that were agreed, and it is not
optional. Answering Paused for a stopped queue claims an intention
where the truth is that nothing is running any more — the same lie L2
removed from the limit getter a commit ago, reintroduced through a new
door. With all four the switch is total and one call answers the
whole question.

The state is computed from fields that already exist; no new
bookkeeping, and the reader count is only read when it has to be. The
order of the two checks is a decision, not a detail, so both orders
are pinned by a test: enabled before paused, because a stopped queue
has no pool to describe; paused before the count, because a limit of
zero means no readers by definition and "somebody set it to zero" is
true where "the pool lost its readers" is a guess.

Side effect worth having: the internal SaWorkNoProgress is gone. The
WaitForIdleAsync warning now prints a value of the public enum, so
there is one enum instead of two and the log cannot drift from the
API. A test asserts they still agree.

Not changed, on purpose: WaitForIdleAsync still decides its reason
with its own checks rather than through PoolState. Unifying them is
tidy and would also reorder the priority, which changes what a
paused-then-stopped queue does when waited on. That is a behaviour
change and wants its own test, not a side effect of observability.

L9 — WaitForIdleAsync took the token first and a flag after it,
against CA1068 and the .NET API design guidelines. With both
parameters optional, a caller who only wants the other one has to name
the token, and every call carries the noise. Now
WaitForIdleAsync(bool failIfNoProgress = false, CancellationToken
cancellationToken = default).

This breaks the binary contract of the published 0.12.0, and the
version has not been bumped — release decision, not mine to make
here. Source-level it is loud rather than silent: a CancellationToken
does not convert to bool, so every positional call is a compile error
instead of a silently changed meaning. All 17 files with call sites
in this repo, Sa.Schedule included, now pass cancellationToken: by
name, which also survives the next reordering.

The guideline is still unenforced: EnableNETAnalyzers is on but
CA1068 is off by default, so the build stays quiet. Turning it on for
the whole solution is a separate decision and was not taken here.

Nine tests in WorkQueuePoolStateTests. The headline one builds two
queues, one force-cancelled and one paused, and asserts they are
indistinguishable by every number the queue used to publish — that is
the finding, stated as code — and different only through PoolState.
Reverting the property to check paused first fails exactly the
precedence test; dropping the _paused check entirely, which is the
distinction itself, fails four.

Three of my own bugs along the way, none of them the product's: a
guessed item count, a paused queue that raced a reader onto its own
item, and a force-cancel wait that expected a buffer the successful
path had already drained.

Sa.Utils.WorkQueue.Tests 143/143 (3 runs), Sa.ScheduleTests 133/133,
solution builds with 0 warnings.
…r language

The feature landed with its contract in the interface and nine tests,
and the readme had never heard of it — zero mentions in both files.
Also missing: the row in the API table, and any mention of emergency
stop in Features, which has been a headline capability since before
any of this work and never got one.

Adds a Pool Observability section to each readme, right after
Concurrency & Scaling because that is where the confusion is created:
the sentence about a reader dying on its own and lowering the limit is
exactly the sentence that leaves a reader of the docs with a limit of
zero and no way to tell a pause from a dead pool. It now says so and
points at the property.

The section leads with why the numbers cannot answer it rather than
listing the enum, since the four names are the easy part and the
reason two of them were previously indistinguishable is the part worth
carrying away. It states the precedence rule for Stopped against
Paused, because a queue paused and then stopped is the case that gets
the answer wrong by accident.

One claim is deliberately weaker than it could be. The section says
that for an enabled queue WaitForIdleAsync names the same two reasons
in its warning, and that a test holds them to that. It does not say
the two cannot disagree: WaitForIdleAsync still decides with its own
checks, so the two would answer differently for a queue that is
stopped. That path is unreachable in practice — a stopped queue is
drained, so the wait returns before it reaches a reason at all — but
"unreachable today" is not the same claim as "impossible", and the
audit records the unification as a separate decision rather than as
something the documentation should promise.

Both files edited in lockstep: 11 feature rows each, the new section
at the same line in both, same table and fence counts.

Sa.Utils.WorkQueue.Tests 143/143, solution builds with 0 warnings.
The warning this method logged, the exception it threw and the state a
caller polls were three answers to one question, kept in agreement by
hand. They now come from a single SaWorkPoolState read.

The one behaviour that changes: a queue that was paused and then stopped
reports Stopped, so the wait parks for the shutdown drain instead of
returning "paused, no progress possible" while that drain is running. It
returned with work still in flight, and IsIdle() said false a moment
later. Under failIfNoProgress it was worse — QueuePaused tells the caller
to raise ConcurrencyLimit, and since L2 that assignment does nothing on a
stopped queue. Wrong advice in an exception costs more than a silent
return.

Parking is safe: the idle signal is released by MarkInactive per item, and
the drain resets the counter only once no reader can still hold one, so the
wait lasts exactly as long as work really is still running. A reader that
outlived the timeout keeps its item counted. This already held for a
stopped queue that was not paused; the change gives the paused case the
same behaviour rather than adding a new way to hang.

Both early returns are untouched and each still names its own cause. HasLive
Readers lost its last caller and is gone.

Three tests in WorkQueueStabilityTests. The proof of "it waited" is a
processor that ignores its cancellation, so a reader outlives the shutdown
and IsIdle() is false after ShutdownAsync returns — without that the
assertion would pass vacuously. Reverting to the old checks fails exactly
the two behaviour tests; the third is a deliberate regression guard and
passes both ways.

The guard test itself was wrong at first: it asserted only
InvalidOperationException, and a mutant making both causes throw
QueueHasNoReaders survived it. Both causes are the same type. It now
asserts the message, and the mutant dies.
Bug fixes:
- AsyncWavReader: fix infinite spin after time trim (cutTo). The reader
  left unconsumed buffer tail and looped in ReadAsync; now stops reading
  right after cutTo. Confirmed: 82M ReadAsync calls in 5s before, ~0 after.
- TimeRange.Seconds: clamp condition was inverted, so any finite toSeconds
  was ignored and [5s, 15s) became [5s, inf). Time-based trimming by
  seconds never worked.
- WavHeaderReader: add WAVE_FORMAT_EXTENSIBLE support (cbSize, validBits,
  channelMask, SubFormat GUID -> PCM/IEEE float mapping), skip extra fmt
  bytes, validate "fmt " after JUNK chunk. New BinaryPipeReader
  ReadInt32Async/ReadGuidAsync helpers; WavHeader gains subformat-aware
  IsPcm/IsIeeeFloat/EffectiveAudioFormat.
- AsyncWavWriter/WavIO: guard >2 GiB RIFF sizes (int32 overflow) with
  InvalidOperationException instead of writing a corrupt header.

Memory/allocations:
- Echo AggressiveMode + WavIO.ReadInterleavedFloatsAsync: reuse
  ArrayPool<float> buffers instead of full ToArray() copies; returned to
  pool after async write completes.
- Echo CrossFeedSeparator.WriteLinearInMemoryAsync: no full .ToArray()
  copy of the file.
- ReadStreamableChunksAsync: internal ConvertToFormatAsync with
  allowBufferReuse: true (channel copy is synchronous before next
  MoveNextAsync), removing one array allocation per sample; per-channel
  trailing offset instead of a shared lastOffset.

Tests: CountingPipeReader (deterministic spin detection), trim mid-file,
extensible PCM WAV, TimeRange.Seconds (finite/open-end), Echo integration
(synthetic stereo WAV with pauses; in-memory, aggressive, streaming-exact).
Sa.MediaTests: 44 total, 39 passed, 5 skipped; solution build 0/0.
…sync

Previously the buffer rented from ArrayPool<float>.Shared was handed to the
caller (AggressiveMode) whose finally-return only ran on the success path.
On a mid-read exception or cancellation the method threw before the tuple
was returned, so the rented buffer leaked out of the pool.

Now the method returns the buffer itself on failure; ownership is only
transferred to the caller on success (docs updated accordingly). Matches
the same-scope Rent/finally-Return pattern used in CrossFeedSeparator and
AsyncWavWriter.
dundich
~
Signed-off-by: dundich <krivchenko-kv@activebt.ru>
…s M1-M9

C2 (expand/contract): __offset$ keeps legacy group_offset UUID plus live
group_offset_seq BIGINT; cursor on msg_seq (SqlLoadConsumerGroup RETURNING
MAX), floor resolvers by date/exclusive key, idempotent ADD COLUMN upgrade
hooks, public API unchanged (WithMinOffset(Guid/DateTimeOffset) preserved).

Optimizations (stage 4):
- M8/Q4: batch ceiling 512 -> 1024 via single DefaultMaxLen const; BatchParams
  links to it; honest param-ratio comments (8+4/N finish, 4 error); pinned by tests
- M1: drop dead ObjectPool<StringBuilder> (policy never retained buffers);
  Build*InsertValues use sized StringBuilder
- M2: one RecyclableMemoryStream per COPY batch, buffer rewound per payload
  (span Write unavailable in Npgsql 10 binary importer - documented)
- M3: TaskQueueReader.Ordinals (14 ints) captured on first row, Get* by index
- M5: error text composed once in GroupByException; ErrorRow carries ErrorMessage
- M6: partition ensure keyed by (day, tenant, part); msg+task and delivery+error
  ensures run in parallel via Task.WhenAll
- M7: ReturnDelivery is a single pass (errs list, delivery parts, error days);
  GetErrors removed
- M9: SqlOutboxBuilder documented as required singleton
- M4: deferred (needs public IOutboxMessageSerializer API extension, R3)

Docs: review statuses (M1-M3, M5-M9 closed, M4 deferred), README updates,
Sa.HybridFileStorage.Postgres package-name typo (Sa.Partial -> Sa.Partitional).

Full solution run: 1355 tests, 0 failed, 5 skipped.
…rn, useless lock

Retries never ran. LoadAsync wrapped every failure in InvalidOperationException
*inside* the delegate handed to PgRetryStrategy, and Retry.WaitAndRetry passes
the escaping exception to shouldRetry without unwrapping InnerException. The
predicate tests `ex is NpgsqlException`, so it always returned false and
rethrew on the first attempt — the advertised "automatic retry" was dead code.
The same catch-all also swallowed OperationCanceledException, so cancellation
could not work either. The wrap now happens after retries are exhausted.

A new IPgDataSource (and therefore a whole connection pool) was created and
disposed on every Load, so every Reload redid TCP, TLS, auth and a Postgres
backend fork. The pool is now built once and reused; ConfigurationProvider is
no longer IDisposable in .NET 10, so the provider implements it directly and
ConfigurationRoot.Dispose reaches it.

The Lock around SetData was pure cost: ConfigurationProvider.TryGet reads Data
without a lock, so guarding the reference assignment never made readers safer.
Replaced by an Interlocked guard that collapses concurrent reloads into one
query.

Also:
- public LoadAsync(CancellationToken); the token is now actually forwarded to
  the query instead of the hardcoded CancellationToken.None
- a failed load keeps the previous snapshot, so a blip cannot wipe the config
- fail-fast validation of options at registration, not at first load
- ToString redacts parameter values and the connection string password
- MaxAttempts / MedianFirstRetryDelay / SkipEmptyValues / LastWins options
- AddSaPostgreSqlConfiguration overload sharing the application's data source
- clearer error when the query does not return two columns

Behaviour notes: a NULL value still cannot shadow a lower-priority source,
because ConfigurationRoot keeps scanning while providers return null — only an
empty string overrides. Documented in both READMEs.
…key behaviour

SkipEmptyValues was documented as "drop rows with a blank key, or a NULL/blank
value", implying it governs the key too. It does not: the blank-key check runs
before the option is consulted, so an empty or whitespace-only key is dropped
unconditionally. An empty key is not a meaningful IConfiguration path and
GetChildKeys misbehaves on one, so the code is right and the docs were wrong.

The asymmetry is now stated in the XML docs and in both READMEs, and pinned by
a test: with SkipEmptyValues = false a blank key is still dropped while a blank
value beside a real key is kept.
…fix README

The sample readme carried three claims the code did not support, and the code
itself demonstrated the worst variant of the new API.

README-only:

- The architecture diagram nested AddSaPostgreSqlConfiguration inside
  AddSaConfiguration(). Not just inaccurate: Sa.Configuration does not reference
  Sa.Configuration.PostgreSql, so the method is not even in scope there. Drawn
  as two independent calls, in both languages.
- The code snippet used app.MapGroup(...) without ever showing
  var app = builder.Build(), so it did not compile.
- "Expected Response" showed User ID=... and an unmasked Password. Npgsql
  normalises the keywords it round-trips, so the real output is Username=... /
  Search Path=..., and the password is now masked. Expected block replaced with
  the verified actual response.

Code:

- select * from settings -> select key, value from settings. The provider reads
  the first two columns as (key, value), so the column order is part of the
  contract; `select *` makes it implicit, and a sample is copied.
- The application already created a data source for seeding, then let the
  configuration source open a second pool for the same database. Handed over ds
  instead, which is what the AddSaPostgreSqlConfiguration(options, IPgDataSource)
  overload exists for. Ownership stays in Program: the provider never disposes a
  data source it did not create.
- /settings returned sa:pg:connection verbatim, i.e. the password, to any caller.
  Masked with NpgsqlConnectionStringBuilder — the value is still shown, since the
  point of the sample is that the secret chain resolved it.

Verified by running the sample against a real database, not just by building:
GET /settings returns Password=*** with theme/language/notifications populated
from the table, which exercises the shared-data-source path end to end.
…ADMEs

The diagram duplicated what the Key Code block already shows and, having been
corrected once, was a liability: it asserted a call-tree relationship that
nobody can verify from the sample, and any future change to the wiring would
silently desync it again. The Configuration Chain section already documents the
secret resolution order in prose.
…l failures

A review of the shared sources in src/Sa (the files that are link-Compile'd
into Sa.Outbox, Sa.Outbox.PostgreSql, Sa.Partial.PostgreSql, Sa.Schedule,
Sa.Data.PostgreSql and Sa.HybridFileStorage.S3) turned up two of each of
the most dangerous kind of bug: a failure that is completely silent, and a
crash that cannot be caught. Five of the findings are fixed here, each with a
regression test.

Every IProcessExecutor timeout was ignored. The CTS was created under
`timeout.HasValue && timeout.Value == TimeSpan.Zero`, so only a *zero*
timeout armed it and any real timeout was discarded: `ExecuteAsync` with a
2s timeout on a 30s process ran the full 30s. Now `t > TimeSpan.Zero`, with
null and TimeSpan.Zero both meaning "no timeout".

Cancellation was ignored too, and did not merely fail to stop early. `Run`
awaited `Task.WhenAll([waitTask, ...readers])`, and WhenAll does not
short-circuit on cancellation - only on fault. The stream readers reach EOF
only when the process closes its pipes, which a hung process never does, so
cancelling either the timeout or the caller's token waited out the process's
natural death: measured 30 002 ms for a 300 ms token. Replaced with
Task.WhenAny, which is what Sa.Media.FFmpeg's own fork of this file already
does. The catch order was wrong as well, so caller cancellation fell into the
catch-all and was reported as ProcessExecutionException; caller cancellation
is now OperationCanceledException and only our own timeout is
ProcessTimeoutException. ExecuteStdOutAsync had the same contract gap and
swallowed its timeout as a bare OCE. Also process.ExitCode -> HasExited ?
ExitCode : -1, because ExitCode throws InvalidOperationException on a live
process, i.e. inside a catch block it would replace the real error.

Two stack overflows, both unrecoverable by definition. NormalizeWhiteSpace
did `stackalloc char[str.Length]`; on a 256 KB-stack thread a 600 002-char
string killed the process (ASP.NET threadpool threads get 1 MB, so ~500k
chars). Levenshtein.CalculateDistance did three `stackalloc int[m]`, i.e. 12
bytes per character of the second string, and 60k characters was enough. Both
now cap the stackalloc and fall back to the heap - ArrayPool for Levenshtein -
matching what GetMurmurHash3 in the same file already did. The loop body moved
to CalculateDistanceCore because a stackalloc result cannot live across a try.

LockRenewer.KeepLocked swallowed a failed renewal into Debug.WriteLine, which
the Release build strips: the loop died after the first failure, the lock
expired, another owner could take it, and the original owner carried on
believing it was still holding the lock. The error is now recorded and thrown
as LockRenewalException from DisposeAsync. Three smaller faults fixed on the
way: the return type was declared IAsyncDisposable, which made the class's own
IDisposable.Dispose() unreachable, so a fire-and-forget caller had no way to
stop renewal; sync Dispose() only disposed the timer and left the loop running
to the next tick; and Task.Run was given the caller's token, so an
already-cancelled token skipped the blockImmediately extension entirely.

Section.IsEmpty() compared the record against `Empty = new(default!, default!)`.
For any value-type T that sentinel is indistinguishable from a legitimately
built point - `new Section<int>(0, 0) == new Section<int>(default, default)` -
so IsEmpty() returned true for a real section. Emptiness is now structural
(Start > End), matching the existing IsPoint/IsPositive vocabulary, and the
three Empty sentinels that caused the collision are gone. They had no callers.

Behaviour note: DisposeAsync now throws when renewal failed, so in
Sa.Outbox/Delivery/DeliveryTenant the `await using` around Deliver will let a
renewal failure mask a delivery failure. A silently expired lock seemed the
worse of the two, but Sa.Outbox has no logger to report it with instead.

Tests: IProcessExecutorTests (10), StackallocTests (8, both overflow repros
on a 256 KB-stack thread plus a differential check across the stack/heap
boundary), SectionTests (12, incl. agreement with InRange), LockRenewerTests
+7. SaTests: 209 total, 209 passed. Solution: 18 assemblies, 1427 total, 0
failed, 5 pre-existing skips; build 0 warnings, 0 errors.
…han once

UseHostedService and AddErrorHandler stored their result in a field on the
IScheduleBuilder instance, but IScheduleSettings was registered with
TryAddSingleton from inside the ScheduleBuilder constructor - so only the
factory from the *first* builder ever ran, and it closed over that first
builder's fields. Every later call was silently discarded:

    services.AddSaSchedule(b => b.AddJob<TestJob>());
    services.AddSaSchedule(b => b.UseHostedService());   // IsHostedService == false

That is the normal way to use this library. Libraries contribute jobs rather
than replacing them - AddSaPartitional registers migration plus cleanup, each
AddDeliveryJob runs per consumer group - so AddSaSchedule is called several
times per container, and only the first call could reach the global settings.
The same trap applied to AddErrorHandler, where a second call replaced the
first handler on its own instance and the surviving settings read a field
nobody else was going to set.

The settings now read what the container holds instead of what one builder
instance happened to remember. UseHostedService registers a
ScheduleHostedServiceMarker (TryAdd, presence is all that matters) and
AddErrorHandler registers an ErrorHandlerRegistration per call;
ScheduleSettings.Create(IServiceProvider) composes the handlers as "any of
them consumed it", which matches the return value's meaning in JobErrorHandler
and, with a single handler, is that handler with no wrapper added.

AddSaSchedule's `configure` is now optional. The settings were registered as a
side effect of `new ScheduleBuilder(...)`, so a call without configure left
IScheduler unresolvable; they are registered in the core path instead, before
the callback, so a caller can still Replace them from inside it.

Also ArgumentNullException guards on both null arguments, and docs for the
now-guaranteed behaviour: handlers accumulate across calls and across
AddSaSchedule calls, UseHostedService is idempotent, every call must complete
before the first IScheduler is resolved (the scheduler snapshots the settings
when constructed), and TimeProvider is registered so a test can replace it
before building the provider.

Tests: ScheduleRegistrationTests (11) - handler and hosted-service flag
surviving a second call, handler composition, all-decline falling through to
the per-job policy, no handler leaving HandleError null, a single host across
repeated calls, jobs accumulating, resolution without configure, both null
guards. Sa.ScheduleTests and the solution build pass.
The project referenced ..\Storage\Storage.csproj, which no longer exists, and
was not listed in Sa.slnx - so it had not been compiled for a long time. It
was also the last thing in the repo still on net8.0 with xunit v2, VSTest,
FluentAssertions 6 and StyleCop, i.e. the exact stack the rest of the
repository has moved off. 762 lines of tests for a sample that is itself gone.

Nothing referenced it: no solution entry, no CI reference, no project
reference from any other sample.
The repo already ships Readme.md in the other packages; these four were the
odd ones out. On a case-insensitive filesystem the rename is invisible, but
git tracks it as R100 and Linux CI builds that only see Readme.md in the
working tree will now agree with the index.

Content untouched - the breaking-changes sections are removed in the next
commit rather than here, so the rename stays reviewable on its own.
Those four packages are all on 0.13.0 now, and v0.13.0 is not tagged yet -
the last tag is v0.12.0. The section documented the migration from a released
version to an unreleased one, so it described breaking changes a consumer
cannot have hit yet, and it would have shipped describing a version that does
not exist. It was written as a pre-release note and belongs in the commit that
ships the release, not in the readme of the package itself.

Also removes the Project Structure section from Sa.HybridFileStorage, whose TOC
entry went with it - the list of internal folders had already drifted away from
the layout. No other section is touched, and the ru/en pairs stay in step.
… into the code

The review landed in 784167a ("~") and its findings have since been worked
through in real commits - d7c585e (C2 offset-seq plus M1-M9) among them - so
the document was a stale scratch file 1051 lines long, describing an audit
that had already been turned into commits and a released change.

The one thing it still held that nothing else did was the manual data step for
converting a *populated* deployment, and SqlOutboxBuilder.cs pointed at it for
exactly that, so deleting the file would have left a dead reference and taken
the instructions with it. Both parts of the step now live where the schema hooks
are defined, next to the ADD COLUMN statements that cannot perform them: the
msg_seq renumber (with the caveat that a plain BIGSERIAL backfill numbers in
physical order, not msg_id order) and the per-group group_offset_seq seed,
without which every cursor starts at 0 and the whole history is delivered
again as new messages.

The hooks keep their promise of only ever adding columns, and the comment
records the run-order dependency between the two statements.
The previous commit made LockRenewerHandle.DisposeAsync throw
LockRenewalException, and left a note that in Sa.Outbox it would let a
renewal failure mask a delivery failure. Traced end to end, that was
optimistic: it broke the cycle, not just the message.

DeliveryTenant.ProcessInTenant holds the renewer in an `await using`, so
the exception is raised while the method unwinds. Nothing between there
and the scheduler absorbs it. ProcessTenantWithTimeout catches only
OperationCanceledException, ProcessTenantsSequential (the default, since
PerTenantMaxDegreeOfParallelism is 1) adds the per-tenant result and keeps
going, the greedy do/while in ProcessMessages breaks, DeliveryJob.Execute
catches nothing, and JobErrorHandler logs the exception and rethrows it.
So if delivery itself had failed, the renewal failure replaced the real
cause; and if delivery and ReturnDelivery had both succeeded, `return
delivered` never ran - the batch was already closed in the database, but
the count was lost and the whole cycle failed. In the parallel pass the
exception came out of the Parallel.ForEachAsync body, which also stopped
every tenant still in flight. That is a regression against the contract
ProcessTenantWithTimeout is written for: it exists so that one
unresponsive tenant cannot fail the whole cycle.

A lost lock is local to its tenant, and so is everything it can cause:
the next pass re-acquires the messages, and the losing owner's batch is
closed by ReturnDelivery like any other. LockRenewalException is now
caught there, next to the cancellation, and counted as 0. Nothing else is
swallowed - a storage failure still fails the cycle, in both passes.

The cost, stated plainly: the count ProcessInTenant never got to return is
lost for that tenant, and the failure is unobserved, because Sa.Outbox has
no logger to report it with. Making it visible needs a logger or a
contract change there, and the per-tenant pass is deliberately where that
boundary sits.

Two smaller things in LockRenewerHandle. DisposeAsync now disposes both
cancellation sources, since the loop is known to have finished by then and
the linked source holds a registration on the caller's token - which for
an application token lives for the lifetime of the process. An interlocked
flag guards it, because the second dispose used to be harmless and would
otherwise call Cancel on a disposed source and report ObjectDisposedException
instead of the lost lock; the failure is still reported on every dispose,
since a caller that swallowed it the first time would otherwise never hear
about it. And an OperationCanceledException arriving while our own token is
not cancelled comes from the extension call itself - Npgsql throws one from
Command.Cancel when the row lock is taken away - so it is recorded with its
own message instead of being lumped in with a generic renewal failure, and
DisposeAsync no longer wraps it a second time.

Tests: DeliveryProcessorTests (6) - a lost lock on one tenant keeps the
counts of the others in both the sequential and the parallel pass, a lost
lock on every tenant reads as an empty cycle, a storage failure still fails
both passes, and a per-tenant timeout is absorbed exactly like a lost lock.
Three of the six fail without the catch. LockRenewerTests +4 - cancellation
from the extension call is reported separately from shutdown, our own
token's cancellation is not reported at all, a repeated DisposeAsync reports
the failure both times, and a sync Dispose after an async one is a no-op.
Sa.Outbox.Tests: 84 total, 84 passed. SaTests: 213 total, 213 passed.
Solution: 18 assemblies, 1437 total, 0 failed, 5 pre-existing skips; build
0 warnings, 0 errors.
Findings 6-21 from the src/Sa review, each with a regression test:

- MergeIntervals([]) returns [] now instead of IndexOutOfRangeException from
  the unconditionally-taken sortedList[0].
- FindEmptyIntervals clips busy intervals to range and ignores those outside
  it, so [10,100] minus [0,5] is [10,100] (was [5,100]) and a busy interval
  entirely to the right no longer yields an empty result.
- GetChunks / GetChunksArray validate chunkSize > 0 up front: a zero size
  stalled enumeration (infinite empty chunks) and GetChunksArray divided by
  zero.
- Retry back-off saturates at TimeSpan.MaxValue via ToDelay instead of
  throwing OverflowException at the 47th exponential delay; a back-off past
  the ceiling means "do not retry within any lifetime".
- Retry.WaitAndRetry preserves the original stack trace with
  ExceptionDispatchInfo instead of `throw lastEx` resetting it to the retry
  frame.
- Retry.IsFatal delegates to ExceptionExtensions.IsCritical; the two lists
  agreed except for AccessViolationException, which made an access violation
  fatal to Retry but retriable to its callers.
- NumericExtensions: the (long) cast of ulong wrapped to a pre-1970 date —
  now validated against long.MaxValue; the double overload floors instead of
  truncating towards zero and rejects NaN/Infinity.
- DateTimeExtensions.ToUnixTimestamp floors (not truncates) so pre-1970
  sub-second stamps no longer round up to the epoch.
- StrToExtensions.StrToDate assumes local uniformly (was a mixed Kind
  depending on which format matched); StrToEnum rejects undefined numeric
  values for non-Flags enums.
- MurmurHash3 loads each block little-endian via BinaryPrimitives instead of
  BitConverter, so the hash is portable across endianness.
- EnumerableExtensions.JoinByString returns string.Empty for a null source
  instead of a null it had silenced with `default!`.
- StringExtensions.NormalizeWhiteSpaceSpan bounds-checks the destination
  buffer; GetMurmurHash3 uses GetByteCount so str.Length*3 cannot overflow.
- WaitForConditionAsync drains the remaining time with a bounded loop that
  crosses the timeout without spinning (the old Task.Delay(remaining) could
  fire early and loop thousands of times).

Regression tests added to SaTests (LockRenewer, MurmurHash3, Retry, Section,
StringExtensions and five new Extensions suites); the full SaTests run is
365 tests, 0 failed.
…ng local

The StrToDate overloads previously OR'd in DateTimeStyles.AssumeLocal before
handing the formats to DateTime.TryParseExact, so an input without an offset
always came back as DateTimeKind.Local and an offset-bearing input was
converted to the host zone. That hid the fact that the same shape of text
meant different things depending on which format matched, and callers could
not choose their own interpretation.

Drop the forced AssumeLocal and pass <paramref name="style"/> through unchanged,
so the returned Kind mirrors what TryParseExact yields: an offset-less string
is Unspecified and a string with an offset is Local, with the moment preserved
(e.g. +03:00 still round-trips to 07:00 UTC). Callers that want a fixed kind
can pass an explicit style such as AssumeUniversal | AdjustToUniversal.

Updated the StrToDate tests to the plain-style contract. SaTests: 365 passed,
0 failed.
…bs to 0.15.0

- README (EN/RU): remove Sa and Sa.Outbox sections; condense remaining
  project descriptions to brief summaries without tables/code
- README (RU): 'Образцы' -> 'Примеры'; drop non-existent Storage.Tests
  sample; add Echo and WorkQueue.Console; fix FFMpeg.Console link casing
- README (EN/RU): trim Sa.Outbox.PostgreSql to the essence (no pathos)
- README (EN/RU): add test-stack note to Architecture section
- all 16 Sa.* libs: set Version 0.15.0
@dundich
dundich merged commit b77a197 into main Sep 29, 2026
0 of 2 checks passed
@dundich
dundich deleted the echo branch October 2, 2026 13:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant