Skip to content

Latest commit

 

History

History
214 lines (151 loc) · 11.8 KB

File metadata and controls

214 lines (151 loc) · 11.8 KB

Testing Guide

Every test mode in this repo, what it exercises, and when to use it. The testing strategy follows the trust boundary: code that parses untrusted bytes gets fuzzed; everything else gets hand-written unit and integration tests.

Test Inventory

Mode What runs Needs container Use when
Unit/integration (go test ./...) All in-process tests: storage engine, broker, server on ephemeral port No Default for development and CI
Race detector Same tests with race detection enabled No Before every commit; CI runs this
Fuzzing (-fuzz=) One fuzz target, generating random input continuously No Probing parsers of untrusted bytes
Fuzz seeds (plain go test) Committed crasher inputs replayed as regression cases No Automatic; part of normal go test
External integration (MQ_ADDR) Create/produce/consume against a live server over TCP Yes Verifying the Docker image end-to-end
Offline validator (mqvalidate) Re-parses every on-disk file: CRC, offsets, index/timeindex consistency No (reads volume) After crashes, upgrades, or before trusting backups

Quick Reference

Unit + integration tests

go test ./...                                  # all packages
go test ./... -race                            # CI configuration
go test ./pkg/log/ -v                          # verbose, one package
go test -run TestServerProduceConsume -v ./pkg/server/   # single test

The server integration tests spin up their own in-process server on a random port, so they never conflict with a running container or each other.

Coverage

go test ./... -coverprofile=coverage.out
go tool cover -func=coverage.out | tail -1     # total %
go tool cover -html=coverage.out               # per-line browser view

CI gates at 77% total coverage. Remaining uncovered lines are file-system error paths and listen failures — impractical to hit without a VFS abstraction layer.

Fuzzing

Four targets exist, all at the serialization layer where bytes cross the trust boundary:

Target Package Property checked
FuzzDecodeMessage pkg/log Arbitrary bytes into DecodeMessage must not panic
FuzzMessageRoundTrip pkg/log Random key/value encode -> decode round-trips exactly
FuzzReadFrame pkg/server Arbitrary bytes into ReadFrame must not panic
FuzzFrameRoundTrip pkg/server Random protobuf payloads frame -> deframe exactly

Run one target at a time (-fuzz accepts only one pattern):

go test ./pkg/log/ -fuzz=FuzzDecodeMessage -fuzztime=30s
go test ./pkg/log/ -fuzz=FuzzMessageRoundTrip -fuzztime=30s
go test ./pkg/server/ -fuzz=FuzzReadFrame -fuzztime=30s
go test ./pkg/server/ -fuzz=FuzzFrameRoundTrip -fuzztime=30s

Notes:

  • Fuzz flags:

    Flag Meaning Default
    -fuzztime Total fuzz duration (30s, 10m) or iterations (1000x) unlimited (until Ctrl-C)
    -fuzzminimizetime Time spent minimizing each crasher found 60s
    -parallel Worker processes GOMAXPROCS
  • How long to run: 30s is fine for regression checks after touching serialization code (~3M execs). Use minutes when changing message format or framing (message.go, protocol.go) — that is where the real bug hid. The seed corpus runs first, so even short runs replay the committed crashers.

  • Fuzzing found a real bug (uint32 underflow + overflow in DecodeMessage) in under 3 seconds that 100%-coverage hand-written tests missed. See PROGRESS.md "Fuzz Testing" learnings.

  • When a fuzzer finds a failure it writes the failing input to testdata/fuzz/<target>/<hash>. These files are committed and replay automatically during every plain go test run, so fixed bugs cannot silently return.

  • Round-trip targets assert decoded == original, catching silent corruption, not just panics.

  • Seeds use empty slices ([]byte{}), never nil — f.Add(nil) fails type checking against []byte parameters.

Re-run a specific historical crasher as a regression test:

go test -run 'FuzzDecodeMessage/2ba3e7ab0de1bada' ./pkg/log/

External integration test (against the Docker container)

docker compose up --build -d                   # server on :9092, data in mq-data volume
MQ_ADDR=localhost:9092 go test -count=1 -run TestExternalServer -v ./pkg/server/
docker compose down                            # volume persists across restarts

Design points:

  • Skipped without MQ_ADDR — plain go test ./... stays green without Docker, keeping CI container-free.
  • -count=1 is required — Go's test cache is keyed on code inputs and environment variables read via os.Getenv, but it cannot see external state. Restarting the container does not invalidate a cached PASS, so without -count=1 you can get an instant ok while never touching the server.
  • Unique topic names per run — topics are named integration-<unixnano>, so repeated runs against the persistent volume never collide.
  • Consume uses the returned partition — produce routes by FNV-1a hash of the key; the response reports which partition received the message. Tests consume from that partition rather than assuming 0.

Offline validator (mqvalidate)

pkg/validate re-parses on-disk files independently of the log package (deliberately not sharing the encoder, so systematic encoding bugs are catchable) and checks:

  • Every .log record: header decodes, CRC matches payload, offsets sequential from base offset, no trailing garbage
  • Every .index entry: points at a valid record at the correct byte position, monotonic
  • Every .timeindex entry: timestamp matches the referenced record, monotonic

Against the Docker volume without stopping the server (read-only):

go build -o mqvalidate ./cmd/mqvalidate/
docker cp mqvalidate message-queue-server-1:/mqvalidate
docker exec message-queue-server-1 /mqvalidate /data   # exit 0 = clean, 1 = problems listed

Or on a local data directory: mqvalidate ./data. Unit tests cover clean data plus injected CRC corruption, offset gaps, wrong index positions, timeindex mismatches, and trailing partial records.

Benchmarks

go test ./pkg/log/ -bench . -benchtime 2s                       # all benchmarks
go test ./pkg/log/ -run xxx -bench BenchmarkPartitionAppend -cpuprofile cpu.out -memprofile mem.out
go tool pprof -top cpu.out                                      # flat view
go tool pprof -http :8080 cpu.out                               # browser flame graph

Coverage: message encode/decode, segment/partition append (including a batch-size sweep), partition read at 10k/1M depth, index binary search at 10/10k/1M entries, and parallel-writer contention. Current findings live in PROGRESS.md "Benchmarks + Profiling".

Profiling methodology

The exact procedure used to produce every performance finding so far (PROGRESS.md "Benchmarks + Profiling"). Reproduce it whenever investigating a benchmark's cost structure:

1. Collect profiles while the target benchmark runs. Profiles only cover the benchmark process, so scope -bench to one benchmark family at a time (-run xxx excludes unit tests):

go test ./pkg/log/ -run xxx \
    -bench 'BenchmarkPartitionAppend' \
    -benchtime 2s \
    -cpuprofile /tmp/opencode/<name>_cpu.out \
    -memprofile /tmp/opencode/<name>_mem.out \
    -blockprofile /tmp/opencode/<name>_block.out

Profile files match the .gitignored *.out pattern; keep them out of the repo and commit findings as text instead.

2. CPU profile — where do cycles go?

go tool pprof -top -nodecount=12 /tmp/opencode/<name>_cpu.out

Read it top-down:

  • A syscall entry (Syscall6, syscall.Syscall) dominating flat% means the path is syscall-bound: count the calls per operation (grep the code path for Write/ReadAt/Sync) and consider batching.
  • Queue-management frames (runtime.*) above ~15% suggest scheduling/GC pressure; check the memory profile before acting.
  • To attribute time to your code rather than leaves: go tool pprof -cum -top ... sorts by cumulative time, and go tool pprof -list 'funcName' shows per-line costs inside one function.

3. Memory profile — who allocates?

go tool pprof -top -sample_index=alloc_space /tmp/opencode/<name>_mem.out

alloc_space totals all allocation; switch to inuse_space if hunting a leak instead of churn. Per-benchmark B/op and allocs/op from the benchmark output itself are often enough — open the profile when you need the culprit function, which is always in the flat column.

4. Block profile — lock contention? Only meaningful with concurrent benchmarks (BenchmarkConcurrentProduce). If mutex wait time is small relative to total runtime, contention is a non-issue — record that conclusion and skip lock-free work. Our Phase 1 finding: ~12% degradation under parallel writers, i.e., not worth optimizing.

5. Interpret against the taxonomy: syscall-bound vs disk-bound vs CPU-bound (see PROGRESS.md). On this machine, no fsync runs on the append path, so "disk-bound" never applies to hot-path benchmarks; dominant syscalls mean batching opportunity, dominant runtime.* means GC/allocation work.

6. Record findings as prose + numbers in PROGRESS.md, citing the benchmark names and percentages — never commit raw profile binaries.

Benchmark workflow (benchstat)

Comparisons are only meaningful against a recorded baseline on the same machine. Every results file starts with an environment header from scripts/bench-env.sh.

One-time setup:

go install golang.org/x/perf/cmd/benchstat@latest

Capturing a baseline (quiesce containers first; -count=10 for statistical significance):

docker compose stop
./scripts/bench-env.sh > bench-results/<name>.env
go test ./pkg/log/ -run xxx -bench . -benchtime 2s -count 10 | tee bench-results/<name>.txt
docker compose start

Comparing a change (-count=5 is usually enough for the comparison side):

go test ./pkg/log/ -run xxx -bench . -benchtime 2s -count 5 | tee bench-results/opt1.txt
~/go/bin/benchstat bench-results/baseline.txt bench-results/opt1.txt

Rules:

  • Baselines use -count=10; per-change runs use -count=5, escalating a single benchmark to 10 if benchstat reports no significant difference where improvement was expected.
  • Always pin -benchtime=2s so every sample amortizes startup jitter identically.
  • Results files are committed to bench-results/ alongside the code state they measure.
  • If benchstat shows "no significant difference" for a change expected to matter, suspect environment noise first: check loadavg in the .env header, rerun before concluding.
  • Benchmark scratch data goes on real disk (~/.cache/mq-bench via benchDir), never /tmp (often RAM-backed tmpfs: misleading numbers, ENOSPC risk). See PROGRESS.md "Batch-Size Sweep" for both pitfalls found the hard way.
  • Write benchmarks must bound file growth (rewind scratch files periodically) or late samples measure a different disk regime than early ones.

CI

.github/workflows/ci.yml runs:

  1. go vet ./...
  2. go test ./... -race with a 77% total coverage gate

CI has no Docker dependency: all tests except TestExternalServer run in-process, and that one self-skips when MQ_ADDR is unset.

Why This Layout

  • In-process servers on ephemeral ports keep unit tests hermetic — no ordering dependencies, no port conflicts, parallelizable with -race.
  • Fuzzing only at the trust boundary concentrates adversarial-input effort where malformed bytes can actually arrive (disk recovery, network peers). Internal state machines get deterministic table-driven tests instead.
  • One env-var-gated test covers the deployment artifact (Docker image) without pulling Docker into CI or duplicating the full test suite against a live server.