osiris_writer: Track diff between writer and replica offsets - #197
osiris_writer: Track diff between writer and replica offsets#197the-mikedavis wants to merge 1 commit into
Conversation
8272b66 to
0fa0df1
Compare
|
In addition to this (or maybe instead of this) we could track replica freshness. The stream coordinator is calculates "freshness" as a requirement in |
0fa0df1 to
680b7a9
Compare
|
I updated this to perform the same calculation as the stream coordinator does: https://github.com/rabbitmq/rabbitmq-server/blob/765d2c5d748f1a3227b97e966a31a73f4b561867/deps/rabbit/src/rabbit_stream_coordinator.erl#L209-L221 (Also see discussion in rabbitmq/rabbitmq-server#15098) So we could use this metric instead of querying the replication state with a call. |
680b7a9 to
82b2ac6
Compare
|
Tick the box to add this pull request to the merge queue (same as
|
This introduces a metric calculated at every batch which records the difference between the timestamps of the last chunks in the logs of the writer and its replicas, and the sum of offsets that need to be replicated. This can be used to watch replicas catch up on replication, or to diagnose situations when replication is being starved out (network-wise) by high-throughput publishing. The replication diff should usually be low but, for streams seeing traffic, non-zero. This calculation is also used in the stream coordinator when adding a member, via `osiris_writer:query_replication_state/1`. With this change the stream coordinator could be updated to use the counter instead of calling the writer.
82b2ac6 to
36a9e40
Compare
|
Tick the box to add this pull request to the merge queue (same as
|
|
I fixed some edge-cases where the gauge would spike during initialization, and I added a sum of offsets that need replication, the "replication backlog". The total value is not useful probably, but seeing it grow or shrink would let you know how replication is doing. |
This introduces a metric calculated at every batch which records the difference between the timestamps of the last chunks in the logs of the writer and its replicas. This can be used to watch replicas catch up on replication, or to diagnose situations when replication is being starved out (network-wise) by high-throughput publishing. The replication diff should usually be low but, for streams seeing traffic, non-zero.
This calculation is also used in the stream coordinator when adding a member, via
osiris_writer:query_replication_state/1. With this change the stream coordinator could be updated to use the counter instead of calling the writer.