Every test mode in this repo, what it exercises, and when to use it. The testing strategy follows the trust boundary: code that parses untrusted bytes gets fuzzed; everything else gets hand-written unit and integration tests.
| Mode | What runs | Needs container | Use when |
|---|---|---|---|
Unit/integration (go test ./...) |
All in-process tests: storage engine, broker, server on ephemeral port | No | Default for development and CI |
| Race detector | Same tests with race detection enabled | No | Before every commit; CI runs this |
Fuzzing (-fuzz=) |
One fuzz target, generating random input continuously | No | Probing parsers of untrusted bytes |
Fuzz seeds (plain go test) |
Committed crasher inputs replayed as regression cases | No | Automatic; part of normal go test |
External integration (MQ_ADDR) |
Create/produce/consume against a live server over TCP | Yes | Verifying the Docker image end-to-end |
Offline validator (mqvalidate) |
Re-parses every on-disk file: CRC, offsets, index/timeindex consistency | No (reads volume) | After crashes, upgrades, or before trusting backups |
go test ./... # all packages
go test ./... -race # CI configuration
go test ./pkg/log/ -v # verbose, one package
go test -run TestServerProduceConsume -v ./pkg/server/ # single testThe server integration tests spin up their own in-process server on a random port, so they never conflict with a running container or each other.
go test ./... -coverprofile=coverage.out
go tool cover -func=coverage.out | tail -1 # total %
go tool cover -html=coverage.out # per-line browser viewCI gates at 77% total coverage. Remaining uncovered lines are file-system error paths and listen failures — impractical to hit without a VFS abstraction layer.
Four targets exist, all at the serialization layer where bytes cross the trust boundary:
| Target | Package | Property checked |
|---|---|---|
FuzzDecodeMessage |
pkg/log | Arbitrary bytes into DecodeMessage must not panic |
FuzzMessageRoundTrip |
pkg/log | Random key/value encode -> decode round-trips exactly |
FuzzReadFrame |
pkg/server | Arbitrary bytes into ReadFrame must not panic |
FuzzFrameRoundTrip |
pkg/server | Random protobuf payloads frame -> deframe exactly |
Run one target at a time (-fuzz accepts only one pattern):
go test ./pkg/log/ -fuzz=FuzzDecodeMessage -fuzztime=30s
go test ./pkg/log/ -fuzz=FuzzMessageRoundTrip -fuzztime=30s
go test ./pkg/server/ -fuzz=FuzzReadFrame -fuzztime=30s
go test ./pkg/server/ -fuzz=FuzzFrameRoundTrip -fuzztime=30sNotes:
-
Fuzz flags:
Flag Meaning Default -fuzztimeTotal fuzz duration ( 30s,10m) or iterations (1000x)unlimited (until Ctrl-C) -fuzzminimizetimeTime spent minimizing each crasher found 60s -parallelWorker processes GOMAXPROCS -
How long to run: 30s is fine for regression checks after touching serialization code (~3M execs). Use minutes when changing message format or framing (
message.go,protocol.go) — that is where the real bug hid. The seed corpus runs first, so even short runs replay the committed crashers. -
Fuzzing found a real bug (uint32 underflow + overflow in
DecodeMessage) in under 3 seconds that 100%-coverage hand-written tests missed. See PROGRESS.md "Fuzz Testing" learnings. -
When a fuzzer finds a failure it writes the failing input to
testdata/fuzz/<target>/<hash>. These files are committed and replay automatically during every plaingo testrun, so fixed bugs cannot silently return. -
Round-trip targets assert decoded == original, catching silent corruption, not just panics.
-
Seeds use empty slices (
[]byte{}), never nil —f.Add(nil)fails type checking against[]byteparameters.
Re-run a specific historical crasher as a regression test:
go test -run 'FuzzDecodeMessage/2ba3e7ab0de1bada' ./pkg/log/docker compose up --build -d # server on :9092, data in mq-data volume
MQ_ADDR=localhost:9092 go test -count=1 -run TestExternalServer -v ./pkg/server/
docker compose down # volume persists across restartsDesign points:
- Skipped without
MQ_ADDR— plaingo test ./...stays green without Docker, keeping CI container-free. -count=1is required — Go's test cache is keyed on code inputs and environment variables read viaos.Getenv, but it cannot see external state. Restarting the container does not invalidate a cached PASS, so without-count=1you can get an instantokwhile never touching the server.- Unique topic names per run — topics are named
integration-<unixnano>, so repeated runs against the persistent volume never collide. - Consume uses the returned partition — produce routes by FNV-1a hash of the key; the response reports which partition received the message. Tests consume from that partition rather than assuming 0.
pkg/validate re-parses on-disk files independently of the log package (deliberately not sharing the encoder, so systematic encoding bugs are catchable) and checks:
- Every
.logrecord: header decodes, CRC matches payload, offsets sequential from base offset, no trailing garbage - Every
.indexentry: points at a valid record at the correct byte position, monotonic - Every
.timeindexentry: timestamp matches the referenced record, monotonic
Against the Docker volume without stopping the server (read-only):
go build -o mqvalidate ./cmd/mqvalidate/
docker cp mqvalidate message-queue-server-1:/mqvalidate
docker exec message-queue-server-1 /mqvalidate /data # exit 0 = clean, 1 = problems listedOr on a local data directory: mqvalidate ./data. Unit tests cover clean data plus injected CRC corruption, offset gaps, wrong index positions, timeindex mismatches, and trailing partial records.
go test ./pkg/log/ -bench . -benchtime 2s # all benchmarks
go test ./pkg/log/ -run xxx -bench BenchmarkPartitionAppend -cpuprofile cpu.out -memprofile mem.out
go tool pprof -top cpu.out # flat view
go tool pprof -http :8080 cpu.out # browser flame graphCoverage: message encode/decode, segment/partition append (including a batch-size sweep), partition read at 10k/1M depth, index binary search at 10/10k/1M entries, and parallel-writer contention. Current findings live in PROGRESS.md "Benchmarks + Profiling".
The exact procedure used to produce every performance finding so far (PROGRESS.md "Benchmarks + Profiling"). Reproduce it whenever investigating a benchmark's cost structure:
1. Collect profiles while the target benchmark runs. Profiles only cover the benchmark process, so scope -bench to one benchmark family at a time (-run xxx excludes unit tests):
go test ./pkg/log/ -run xxx \
-bench 'BenchmarkPartitionAppend' \
-benchtime 2s \
-cpuprofile /tmp/opencode/<name>_cpu.out \
-memprofile /tmp/opencode/<name>_mem.out \
-blockprofile /tmp/opencode/<name>_block.outProfile files match the .gitignored *.out pattern; keep them out of the repo and commit findings as text instead.
2. CPU profile — where do cycles go?
go tool pprof -top -nodecount=12 /tmp/opencode/<name>_cpu.outRead it top-down:
- A syscall entry (
Syscall6,syscall.Syscall) dominating flat% means the path is syscall-bound: count the calls per operation (grep the code path forWrite/ReadAt/Sync) and consider batching. - Queue-management frames (
runtime.*) above ~15% suggest scheduling/GC pressure; check the memory profile before acting. - To attribute time to your code rather than leaves:
go tool pprof -cum -top ...sorts by cumulative time, andgo tool pprof -list 'funcName'shows per-line costs inside one function.
3. Memory profile — who allocates?
go tool pprof -top -sample_index=alloc_space /tmp/opencode/<name>_mem.outalloc_space totals all allocation; switch to inuse_space if hunting a leak instead of churn. Per-benchmark B/op and allocs/op from the benchmark output itself are often enough — open the profile when you need the culprit function, which is always in the flat column.
4. Block profile — lock contention? Only meaningful with concurrent benchmarks (BenchmarkConcurrentProduce). If mutex wait time is small relative to total runtime, contention is a non-issue — record that conclusion and skip lock-free work. Our Phase 1 finding: ~12% degradation under parallel writers, i.e., not worth optimizing.
5. Interpret against the taxonomy: syscall-bound vs disk-bound vs CPU-bound (see PROGRESS.md). On this machine, no fsync runs on the append path, so "disk-bound" never applies to hot-path benchmarks; dominant syscalls mean batching opportunity, dominant runtime.* means GC/allocation work.
6. Record findings as prose + numbers in PROGRESS.md, citing the benchmark names and percentages — never commit raw profile binaries.
Comparisons are only meaningful against a recorded baseline on the same machine. Every results file starts with an environment header from scripts/bench-env.sh.
One-time setup:
go install golang.org/x/perf/cmd/benchstat@latestCapturing a baseline (quiesce containers first; -count=10 for statistical significance):
docker compose stop
./scripts/bench-env.sh > bench-results/<name>.env
go test ./pkg/log/ -run xxx -bench . -benchtime 2s -count 10 | tee bench-results/<name>.txt
docker compose startComparing a change (-count=5 is usually enough for the comparison side):
go test ./pkg/log/ -run xxx -bench . -benchtime 2s -count 5 | tee bench-results/opt1.txt
~/go/bin/benchstat bench-results/baseline.txt bench-results/opt1.txtRules:
- Baselines use
-count=10; per-change runs use-count=5, escalating a single benchmark to 10 if benchstat reports no significant difference where improvement was expected. - Always pin
-benchtime=2sso every sample amortizes startup jitter identically. - Results files are committed to
bench-results/alongside the code state they measure. - If benchstat shows "no significant difference" for a change expected to matter, suspect environment noise first: check
loadavgin the.envheader, rerun before concluding. - Benchmark scratch data goes on real disk (
~/.cache/mq-benchviabenchDir), never/tmp(often RAM-backed tmpfs: misleading numbers, ENOSPC risk). See PROGRESS.md "Batch-Size Sweep" for both pitfalls found the hard way. - Write benchmarks must bound file growth (rewind scratch files periodically) or late samples measure a different disk regime than early ones.
.github/workflows/ci.yml runs:
go vet ./...go test ./... -racewith a 77% total coverage gate
CI has no Docker dependency: all tests except TestExternalServer run in-process, and that one self-skips when MQ_ADDR is unset.
- In-process servers on ephemeral ports keep unit tests hermetic — no ordering dependencies, no port conflicts, parallelizable with
-race. - Fuzzing only at the trust boundary concentrates adversarial-input effort where malformed bytes can actually arrive (disk recovery, network peers). Internal state machines get deterministic table-driven tests instead.
- One env-var-gated test covers the deployment artifact (Docker image) without pulling Docker into CI or duplicating the full test suite against a live server.