From 911360c5fd7340c95c4f7e73f7be3add419dea40 Mon Sep 17 00:00:00 2001 From: dirkjink Date: Mon, 7 Sep 2026 22:21:39 +0200 Subject: [PATCH] Align the guides with the source code MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The source is authoritative; every claim in the operator guides was checked against the current classes and the drift fixed: - Configuration guide: TOC entry for section 13; the full property composition (conditional acks default, mandatory layer, max.block.ms cap, idempotence validation) in §4 and Appendix A.4; fixed key source and unreadable-MDC behavior in §6; the throttle gap is fixed in code; every fallback diversion reason, fallback ownership and the no-SLF4J-logging rule in §7; the ordered stop() sequence with all budgets, the ~18 s worst case and the no-restart rule (ADR-0004) in §8; bound-state binding decision, rebind semantics and the Logback reconfiguration gap in §9; corrected events.dispatched and fallback.dropped semantics; sendQueueCapacity, mapping and idempotence checks plus the cap warning and runtime worker-death warnings in §10; the send-dispatcher row now states drain, grace and shared budget, and two new "code" rows (key length, max.block.ms cap); the A.2 flow shows the send dispatcher step. - README: required-element wording, all fallback triggers, BlockHound note reflects the send workers, fallback.dropped causes, a working non-Spring binding sample with rebind semantics, and a truthful build section (JDK 24+, quality gates via CONTRIBUTING). - Metrics overview: accepted/dropped/send.error/shutdown semantics, appender tag on the Kafka bridge, alert wording. - Example XML: parser whitespace rule, fallback triggers, why no AsyncAppender. - benchmarks/README: RecordHeadersBenchmark is historical evidence, SenderPathBenchmark doubles as the scalar-replaceability guard, all results folders named. - CONTRIBUTING: DocumentationContractTest pins the tables. - KafkaAppender.addAppender KDoc no longer promises a restart. mvn -o verify (DocumentationContractTest, ktlint) and dokka:dokka pass. Co-Authored-By: Claude Fable 5.1 Claude-Session: https://claude.ai/code/session_015GfdGp7eUrjJKBvcJUx3q2 --- CONTRIBUTING.md | 5 +- README.md | 52 +++-- benchmarks/README.md | 9 +- docs/config/example-logback-spring.xml | 17 +- docs/config/kafka-appender-config-guide.md | 205 ++++++++++++++---- docs/metrics/metrics-overview.md | 14 +- .../eu/inqudium/tabellarium/KafkaAppender.kt | 6 +- 7 files changed, 220 insertions(+), 88 deletions(-) diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index b6f333d..63640fd 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -134,7 +134,10 @@ JAZZER_FUZZ=1 mvn -Dtest=TopicRouterFuzzTest test 2. Keep the change focused — one logical change per pull request. 3. Make sure `mvn verify` passes. 4. Update the README / `docs/` when the configuration surface or the - metrics inventory changes. + metrics inventory changes. `DocumentationContractTest` pins the + configuration guide's defaults tables and the metrics overview's + inventory to the constants, so a changed constant fails `mvn verify` + until the tables follow. 5. Open the pull request with a description of *what* changed and *why*. ## Reporting bugs and requesting features diff --git a/README.md b/README.md index 81403a1..64b14bc 100644 --- a/README.md +++ b/README.md @@ -193,9 +193,11 @@ Reference for every supported element: | `` | No | boolean | Captures caller data on the logging thread before the asynchronous hand-off (default false); only relevant when a fallback layout uses `%caller`. | | `` | No | ref | Single fallback appender — see [Resilience](#resilience). | -Missing or blank values for the five required elements cause the -appender to refuse startup with an explicit `addError` on Logback's -status manager. The error message identifies which element is missing. +A missing `` or a blank ``, `` or +`` makes the appender refuse startup with an explicit `addError` +on Logback's status manager naming the element; a missing +`` or unusable producer properties fail the pipeline +construction the same way. ## Delivery guarantees @@ -302,10 +304,12 @@ Three resilience mechanisms run independently per topic class: source of truth for delivery outcome, so delivery failures are never invisible. -3. **Fallback appender.** When the circuit is open or a send fails - synchronously, the original `ILoggingEvent` is routed to the - configured fallback appender. Standard Logback `` - syntax is supported: +3. **Fallback appender.** Whenever an event cannot reach Kafka — an + open breaker, a throttled probe, a failed send (synchronous or via + the callback), a full send queue, a hot-path error, or the remainder + at shutdown — the original `ILoggingEvent` is routed to the + configured fallback appender, tagged with the reason in the metrics. + Standard Logback `` syntax is supported: ```xml @@ -443,7 +447,7 @@ hazard — `synchronized` blocks in the appender hot path, which cause carrier-thread pinning on virtual threads and Reactor-Netty event-loop stalls — does not arise: the appender extends `UnsynchronizedAppenderBase`. There are no locks in the hot path; -only atomics and volatiles. +only atomics, volatiles and a per-thread reentry flag. Two reactive-specific concerns remain that are worth tuning per service. @@ -505,12 +509,12 @@ buffered record. Services that run BlockHound (`io.projectreactor.tools:blockhound`) in their integration tests will see the appender's internal operations flagged as blocking — most notably the -`LinkedBlockingQueue.offer()` in the [FallbackDispatcher] and the -internals of `KafkaProducer.send()`. These are not true blocks in -the harmful sense (the queue offer is non-blocking on a non-full -queue; the producer send is the operator's accepted -`max.block.ms` budget), but BlockHound's heuristics don't know -that. +`LinkedBlockingQueue.offer()` of the per-class send queue (every +event) and of the fallback dispatcher. These are not true blocks in +the harmful sense (the offer is non-blocking, a full queue rejects +instead of waiting; `KafkaProducer.send()` itself runs on the +appender's own worker threads, never on the caller), but BlockHound's +heuristics don't know that. Add an allow-list entry in the test setup: @@ -594,7 +598,7 @@ Micrometer on the classpath and emits no metrics until | `kafka.appender.events.dispatched` | Counter | `topic.class` | Events handed to `producer.send` without a synchronous failure (callback outcome unknown) | | `kafka.appender.events.fallback` | Counter | `topic.class`, `reason` | Events diverted from Kafka (to the fallback if configured, otherwise dropped) | | `kafka.appender.send.duration` | Timer | `topic.class`, `outcome` | Wall-clock send duration from invocation to callback | -| `kafka.appender.fallback.dropped` | Counter | — | Events lost because the fallback dispatcher queue was full | +| `kafka.appender.fallback.dropped` | Counter | — | Events lost by the fallback dispatcher (queue full, `doAppend` threw, worker died, shutdown remainder) | | `kafka.appender.fallback.queue.size` | Gauge | — | Current depth of the fallback dispatcher queue | | `kafka.appender.fallback.queue.capacity` | Gauge | — | Maximum depth of the fallback dispatcher queue | | `kafka.appender.send.queue.size` | Gauge | `topic.class` | Current depth of the class's send dispatcher queue | @@ -692,15 +696,17 @@ happens after the `MeterRegistry` is available: ```kotlin val loggerContext = LoggerFactory.getILoggerFactory() as LoggerContext loggerContext.loggerList.asSequence() - .flatMap { logger -> - generateSequence({ logger.iteratorForAppenders() }) { null } - .first().asSequence() - } + .flatMap { logger -> logger.iteratorForAppenders().asSequence() } .filterIsInstance() .distinct() .forEach { it.bindMeterRegistry(meterRegistry, Tags.empty()) } ``` +(Descend into `AsyncAppender` wrappers yourself if you use them; the +Spring binding does.) A repeated `bindMeterRegistry` call replaces the +previous binding, and a call on a stopped appender is ignored with a +status warning. + Pre-Spring log events (Logback initialization, Spring bootstrap logging) are not counted in either setup — this is a deliberate trade-off, since capturing them would require a static @@ -861,8 +867,12 @@ AUDIT record end-to-end against an Apache Kafka container — real serializers, compression, headers, and the AUDIT acks/idempotence handshake. -The module has no Maven plugins beyond the Kotlin compiler and Surefire. -Java 21 and Kotlin 2.4.10. +The artifact targets Java 21 (Kotlin 2.4.10); building needs JDK 24+ +because of the JVM flags in `.mvn/jvm.config`. The quality gates that +run with `mvn verify` (ktlint, JaCoCo, the documentation-contract test +that pins the guide's tables to the constants, Jazzer regression +inputs) and the CI-only scans (OSV via CycloneDX SBOM, CodeQL, nightly +fuzzing) are described in [CONTRIBUTING.md](CONTRIBUTING.md). ## Contributing diff --git a/benchmarks/README.md b/benchmarks/README.md index 9be0c7b..0b41fd6 100644 --- a/benchmarks/README.md +++ b/benchmarks/README.md @@ -24,15 +24,16 @@ java -jar benchmarks/target/benchmarks.jar AppendPipelineBenchmark -bm sample -t ``` All benchmarks use 3 forks, 5 warmup and 5 measurement iterations by -default (annotation-driven). Raw outputs of the 2026-08-29 verification -session live under `results/2026-08-29/`. +default (annotation-driven). Raw outputs live under `results//`: +the 2026-08-29 verification session, the 2026-08-30 re-run after the +header fix, and the 2026-09-07 sender-path re-measurement. ## Inventory | Benchmark | Verifies | What it measures | |---|---|---| -| `RecordHeadersBenchmark` | PERF_ANALYSIS-2026-08-29T11-01-08 finding 2 | Per-record header construction: production shape (5 × `headers().add`) vs. one shared pre-built header list. Primary metric: `gc.alloc.rate.norm`. | -| `SenderPathBenchmark` | finding 3 | The delivered-path worker side (`ResilientMessageSender.send` incl. breaker, record build, instant-success callback) with `metricsBound=false/true`; the param delta is the observability envelope per delivered event. | +| `RecordHeadersBenchmark` | PERF_ANALYSIS-2026-08-29T11-01-08 finding 2 | Per-record header construction: the pre-fix shape (5 × `headers().add` per record) vs. the shared pre-built header list production uses since the fix (`EnrichedRecord.headers`). Historical evidence, not a current-code regression guard. Primary metric: `gc.alloc.rate.norm`. | +| `SenderPathBenchmark` | finding 3 | The delivered-path worker side (`ResilientMessageSender.send` incl. breaker, record build, instant-success callback object) with `metricsBound=false/true`; the param delta is the observability envelope per delivered event. Also the guard for the callback object's scalar replaceability: 112 B/op unbound since the `@Volatile` removal (`results/2026-09-07/`). | | `HandoffBenchmark` | finding 1 | The caller-side hand-off primitive: N producers `offer` into a bounded `LinkedBlockingQueue` while one batch-draining consumer keeps it near-empty (worst-case put-lock contention; `rejected` aux counter proves the regime). Thread split via `-tg 1,`. | | `AppendPipelineBenchmark` | secondary evidence (findings 1/3), caller-side allocation | The full production `doAppend` path open-loop against a `DiscardingProducer`. Under open load it saturates the worker and measures the shedding regime — the teardown prints the achieved path mix (delivered vs. diverted), which is part of the evidence, not a hidden variable. | diff --git a/docs/config/example-logback-spring.xml b/docs/config/example-logback-spring.xml index 30af221..7a2590b 100644 --- a/docs/config/example-logback-spring.xml +++ b/docs/config/example-logback-spring.xml @@ -83,7 +83,8 @@ + comments. Leading whitespace (XML indentation) is ignored, + trailing whitespace on values is trimmed. --> bootstrap.servers=${KAFKA_BOOTSTRAP_SERVERS:-kafka.example.com:9092} security.protocol=SSL @@ -155,9 +156,10 @@ Recommended: leave unset or set to false. --> - + @@ -165,9 +167,10 @@