ostia is a profiling and benchmarking toolkit for Bun. It times subprocess commands
(like hyperfine) and in-process functions (like mitata), optionally captures CPU
profiles, heap snapshots, JIT tiers, retained heap and peak memory, and writes everything
to one schema-versioned JSON document (ProfileDocument). Two documents compare with a
bootstrap confidence interval and a Mann-Whitney test, with the regression threshold
widened to the machine's measured noise floor. ostia ab pairs the working tree against
a git ref in one process, which holds up on machines too noisy for that. ostia ci gates
a config file of workloads against a saved baseline, and --format minimal gives scripts
and LLM agents a compact JSON line protocol. The CLI is a thin wrapper over the library, so anything
ostia time/ostia bench do, time()/bench() do too.
Zero runtime dependencies. Requires Bun ≥ 1.4.
bun add ostiaostia time --samples 10 "bun fixtures/fast.ts" "bun fixtures/slow.ts"Apple M2 · 8 cores · load 3.6 · noise floor 0.5%
Task Median Spread Range User/Sys Relative
---------------------------------------------------------------------------------------------------
bun fixtures/fast.ts 9.38 ms 9.52 ms…9.74 ms 9.10 ms…9.76 ms 6.98 ms/3.06 ms 1.00×
bun fixtures/slow.ts 23.2 ms 23.3 ms…23.8 ms 23.0 ms…23.9 ms 20.9 ms/2.89 ms 2.47× slower
! outliers-detected
Warnings:
bun fixtures/slow.ts: 1 outlier(s) detected (1 severe, 0 mild).
The header line shows the machine, its load average, and the noise floor from a ~200ms
reference measurement taken once per run (--no-noise-check skips it). Spread is
p75…p99; User/Sys is the median user/system CPU time per trial.
// suite.ts
import { group, task } from "ostia"
const input = Array.from({ length: 2_000 }, (_, i) => i % 500)
group("dedupe", () => {
task("naive (indexOf scan, O(n²))", () => dedupeNaive(input))
task("Set-based (O(n))", () => [...new Set(input)])
})ostia bench suite.tsApple M2 · 8 cores · load 3.5 · noise floor 0.4%
Task Median Spread Range Relative
------------------------------------------------------------------------------------------------
dedupe:
dedupe/naive (indexOf scan, O(n²)) 173.9 µs 180.4 µs…205.8 µs 170.1 µs…635.5 µs 7.40× slower
dedupe/Set-based (O(n)) 23.5 µs 24.9 µs…72.5 µs 18.7 µs…208.3 µs 1.00×
Commit the suite, then change the code: say the Set-based dedupe becomes
input.filter((x, i) => input.indexOf(x) === i).
ostia ab suite.ts # every task: working tree vs HEAD, paired in one processA/B: working tree vs HEAD (3945768) · 15 rounds · threshold 10% · geomean threshold 1.5%
Task Base Candidate Change p25…p75 Verdict
-------------------------------------------------------------------------------------------------
dedupe:
dedupe/naive (indexOf scan, O(n²)) 183.5 µs 178.5 µs -1.9% -3.5%…-0.3%
dedupe/Set-based (O(n)) 25.2 µs 181.2 µs +618.1% +554.2%…+659.2% regressed, confirmed (repeats: +630.6%, +584.5%)
Geomean +165.4% (threshold 1.5%) · 1 regressed, 0 improved, 1 unchanged of 2 · fail
Base and candidate run in alternating ~10ms batches, so machine drift cancels within each round; a flagged task counts only if two fresh processes agree. Exit 1 on a regression.
// ostia.config.json
{
"samples": 10,
"workloads": [
{ "label": "work", "command": ["bun", "fixtures/work.ts"], "inputs": ["fixtures/**"] }
]
}ostia baseline save # on known-good code: writes .ostia/baselines/main.json
ostia ci # on your change: exit 1 on a regression1 workloads
0 cached
1 executed
0 passed 1 regressed (+44.1% median on work)
Profile CI: ✗
...
✗ work
timing: +44.1% median, 95% CI [+41.4%, +45.6%], p<0.001 (regressed)
Scratch output (cache, artifacts) goes to node_modules/.cache/ostia. Baselines go to
.ostia/baselines/ so they survive reinstalls; add .ostia/ to .gitignore.
--format minimal (on time, bench, ab, compare, report, ci) prints one JSON
object per line on stdout and nothing else. Every line has event and
protocolVersion: 1. Timing values are in nanoseconds.
ostia time --samples 10 "bun a.ts" --format minimal
ostia compare before.json after.json --format minimal
ostia ci --format minimal; echo $?event |
When | Key fields |
|---|---|---|
run |
One per timing measurement, every command | workloadId, task, group?, params?, skipped?, unit, samples, batch, mean/median/stddev/stddevPct/min/max/p75/p99/mad, userNs/systemNs (subprocess only), retainedBytesPerOp?/peakBytes? (--alloc/--peak-mem), relative?, noiseFloorPct?, warnings[], on compare/ci: delta: { medianPct, meanPct, verdict, pass, ci95?, pValue?, effectiveTimingPct, matched }, and on ab: paired: { baseMedian, medianRatio, p25, p75, rounds, verdict, flagged?, confirmed?, repeats?, sameOutput } |
unmatched |
One per workload on only one side of compare/ci/ab |
workloadId, task, side: "base" | "cand" |
summary |
Last line of compare/ci/ab only |
command, matched/regressed/improved/unchanged/unmatched, cached/executed/failed/missingBaseline (ci), geomeanPct, effectiveTimingPct, noiseFloorPct?, baseline? (ci), base?/geomeanThresholdPct?/unconfirmed?/outputDiffers? (ab), git?, exportedTo?, verdict, exitCode |
{"event":"run","protocolVersion":1,"schemaVersion":2,"workloadId":"wl_11e8562f3622d528","task":"work","unit":"ns","samples":10,"batch":1,"mean":21012800,"median":20999900,"stddev":231456,"stddevPct":1.1015,"min":20664000,"max":21552300,"warnings":[{"code":"outliers-detected","data":{"mild":1,"severe":0}}],"p75":21086100,"p99":21517600,"mad":126625,"userNs":15519000,"systemNs":6015500,"noiseFloorPct":2.09286,"delta":{"medianPct":44.0989,"meanPct":43.9626,"verdict":"regressed","pass":false,"effectiveTimingPct":10,"matched":true,"ci95":[41.4394,45.5841],"pValue":0.000157103}}
{"event":"summary","protocolVersion":1,"command":"ci","matched":1,"regressed":1,"improved":0,"unchanged":0,"unmatched":0,"geomeanPct":44.098920968212305,"effectiveTimingPct":10,"verdict":"fail","exitCode":1,"cached":1,"executed":0,"failed":0,"missingBaseline":0,"baseline":{"name":"main","path":".ostia/baselines/main.json"},"noiseFloorPct":2.09286}
Within protocolVersion: 1, keys are only ever added, never renamed or removed.
Exit codes, the same for every command:
| Code | Meaning |
|---|---|
0 |
Pass |
1 |
At least one workload regressed (compare/ci/ab only; time/bench never return 1) |
2 |
Harness error: a command exited non-zero or produced no samples, a suite failed, nothing matched, a bad flag, a missing/invalid config or baseline |
130 |
Cancelled with Ctrl-C (time/bench/ab; partial results are still exported) |
On exit 2, stderr's last line is {"event":"error","protocolVersion":1,"code":...,"message":...,"data"?:...}
when stderr is not a TTY or a machine format (minimal/json/jsonl) was requested.
A person at a terminal sees only the prose message. code is one of invalid-flag,
config-missing, config-invalid, baseline-missing, no-matches, spawn-failed,
command-failed, timeout, time-source-no-match, document-load-failed,
no-cpu-evidence, internal. Full reference: docs/agent-protocol.md.
Every command takes --help. Per-flag detail is in docs/cli.md.
Times commands as subprocesses. --cpu/--heap add one separate instrumented trial each;
the profiler never runs during timing trials.
ostia time "bun a.ts" "bun b.ts"
ostia time --samples 25 --cpu --heap "bun src/server.ts"
ostia time --prepare "rm -rf dist" "bun build.ts"
ostia time --time-source "built in (\d+)ms" "bun build.ts"- Each command string is whitespace-split into argv, with no shell. Everything after
--is one more command's argv, verbatim:ostia time -- bun -e "console.log('a b')". - Default sampling: 3 warmup trials, then trials until ~3s have elapsed and at least 10
ran.
--samples Ngives an exact count per command;--budget MS/--min-samples Ntune the loop. - With 2+ commands, trials round-robin across commands (
--no-interleaveto run them one after another). - A command stops at its first non-ignored non-zero exit, and
ostia timeexits 2.--ignore-failure[=CODE,...]treats the listed codes (bare: all) as success.
Runs in-process group()/task() suites. Each suite file runs in its own child process.
ostia bench bench/*.ts
ostia bench bench/*.ts --filter parse --cpu --alloc
ostia bench bench/*.ts --filter large --peak-mem
ostia bench bench/*.ts --isolate
ostia bench --preload ./bench/dom-setup.ts --bun-flags="--conditions=browser" bench/*.ts- Each task samples for
--budgetms (default 500). Fast calls are batched so one trial spans at least 1µs; the budget-driven loop stops at 20,000 trials. --isolateruns every task in its own process, isolating JIT state, builtin call-site feedback (e.g.Array.prototype.map) and GC heap from other tasks. Use it when you need the most comparable numbers.--jobs N|autoruns suite files in parallel. Faster, but noisier; keep the default of 1 for anything youcompareor gate inci.--cpuprofiles each task at 100µs for about 2,000 samples (--cpu-intervalchanges the interval). Inlined helpers count as their callers' self time.--allocreports the heap each call retains after a full GC: a leak check, not an allocation count.--peak-memreports how far the task's first call raises RSS, garbage included, in 3 fresh processes (OSTIA_PEAK_MEM=1is set there, so a suite can skip heavy setup that would peak first).- With no files,
ostia benchuses the config'sbenchsection. Each flag overrides its config field;--no-gc/--no-cpu/--no-alloc/--no-peak-mem/--no-isolateoverride a configtrue.
Runs suites on a git ref's committed tree and on the working tree in one process, alternating short batches, and gates on the per-round time ratio.
ostia ab bench/*.ts # vs HEAD
ostia ab bench/*.ts --base origin/main
ostia ab bench/*.ts --threshold 5 --rounds 21- The same suite files as
ostia bench, unchanged. The ref's tree is extracted once per commit undernode_modules/.cache/ostia/ab/; relative imports resolve within each tree, package imports to the project'snode_modules. - A task is flagged when its median ratio moves past
--threshold(default 10%) in at least three quarters of rounds, and counts only if--confirm(default 2) fresh processes agree. The run also fails when the geometric mean of all ratios is more than--geomean-threshold(default 1.5%) slower. - Tasks whose first call returns different values on each side are listed (not a
failure). Exit:
0pass,1regression,2nothing paired or a harness error.
Matches two documents' workloads by id and reports a verdict per workload.
ostia compare before.json after.json
ostia compare after.json --baseline .ostia/baselines/main.json
ostia compare before.json after.json --format markdown✗ bun fixtures/work.ts
timing: +23.8% median, 95% CI [+18.3%, +30.1%], p<0.001 (regressed)
A regression needs the whole 95% CI above the threshold and a Mann-Whitney p-value below
alpha (default 0.01). Thresholds come from the config file when one exists, otherwise
the defaults (timingPct: 5). The bootstrap is seeded from the samples, so the same two
documents always give the same verdict. See docs/statistics.md.
Exit: 0 pass, 1 regression, 2 nothing matched or a load error.
Renders a saved document without re-running anything.
ostia report doc.json --format markdown
ostia report doc.json --format minimal
ostia report doc.json --format speedscope --out-dir viz/
ostia report doc.json --format collapsed | flamegraph.pl > flame.svgFormats: table (default), json, jsonl, markdown, minimal, and, for documents
with CPU evidence, collapsed, mermaid, speedscope, cpuprofile. time, bench,
compare and ci accept only the first five; export a document and use report for the
visualization formats.
Runs the config's workloads, compares them against a named baseline, and exits 1 on a regression.
ostia ci
ostia ci --full # ignore the cache
ostia ci --baseline release
ostia ci --save-baseline # after a pass, make this run the new baseline- Command workloads are cached by their declared
inputs: noinputsfield always reruns;inputs: []means "depends on nothing" and caches; otherwise the run is reused while the matched files' contents are unchanged.suitesworkloads always run. suitesworkloads run with the config'sbenchsection, the same wayostia benchreads it.- Exit 2 if any command workload exits non-zero (not ignored) or produces no samples, or
if the baseline file is missing. A baseline that matches none of the configured
workloads is also an error; one missing only some lists them and carries on
(
onMissingBaselinein the config changes this).
ostia baseline save # measure the configured workloads -> .ostia/baselines/main.json
ostia baseline save my-feature
ostia baseline list
ostia baseline show main --format markdownsave uses the same measurement code path as ci. show accepts report's flags.
import {
time, bench, ab, group, task, sweep, range, run, profile, keep,
compareDocuments, defineConfig, createDocument, loadDocument, saveDocument, renderers,
} from "ostia"
import type { ProfileDocument, MinimalEvent } from "ostia"| Export | Does |
|---|---|
time(opts) |
Subprocess timing, same as ostia time. Returns a ProfileDocument. |
bench(opts) |
Runs suite files, same as ostia bench. |
ab(opts) |
Paired A/B of suite files against a git ref, same as ostia ab. |
group(name, fn, opts?) / task(name, fn, opts?) |
Register in-process tasks; .skip/.only variants. |
sweep(dims, fn) / range(start, end, mult?) |
Parameter sweeps; tasks inherit the point as params. |
run(opts?) |
Runs the tasks registered in the current file, in this process (bun suite.ts). |
profile(fn, opts?) |
In-process CPU capture; origin: "jsc" adds JIT tier data. |
keep(value) |
Pins an intermediate value against dead-code elimination. |
compareDocuments(base, cand, thresholds?) |
Same comparison as ostia compare. |
defineConfig(config) |
Typing helper for ostia.config.ts. |
createDocument / loadDocument / saveDocument |
Build, read (schema v2 only), and write documents. |
renderers |
table, markdown, json, jsonl, minimal, collapsed, mermaid, speedscope, cpuprofile. |
const doc = await time({
commands: ["bun a.ts", { command: "bun b.ts", label: "b", prepare: "rm -rf dist" }],
samples: 20,
cpu: true,
})
group("parse", () => {
sweep({ size: range(100, 10_000) }, ({ size }) => {
const input = buildInput(size) // unmeasured setup, once per point
task("parse", () => parse(input), { isolate: true })
})
})
const result = compareDocuments(await loadDocument("before.json"), doc)
if (result.summary.verdict === "fail") process.exitCode = 1
const { text } = await renderers.markdown.render(doc, {})Full reference, including task options, hooks, and a mitata/hyperfine migration table: docs/library.md.
ostia.config.ts (checked first) or ostia.config.json, in the current directory.
// ostia.config.ts
import { defineConfig } from "ostia"
export default defineConfig({
baseline: "main",
samples: 15, // command workloads; or budgetMs/minSamples
warmup: 3,
thresholds: { timingPct: 5 },
workloads: [
{ label: "cold-start", command: ["bun", "src/cli.ts", "--help"], inputs: ["src/**/*.ts"] },
{ label: "spawn", command: ["bun", "-e", "1"], inputs: [] },
{ label: "build:cold", command: ["bun", "build.ts"], prepare: "rm -rf dist" },
{ label: "suites", suites: ["bench/*.ts"] },
],
bench: { budgetMs: 500, isolate: true, preload: ["bench/setup.ts"] },
})A config that still uses the old runs field fails to load with a message naming
samples (error code config-invalid). All fields: docs/config.md.
- docs/cli.md: every command and flag
- docs/config.md: config file reference
- docs/library.md: library API reference
- docs/agent-protocol.md:
--format minimal, exit codes, error codes - docs/statistics.md: sampling, the comparison test, noise floor
- docs/document-schema.md:
ProfileDocument, workload ids, warnings - docs/preload-recipes.md: jsdom, happy-dom and
Bun.plugin()preloads
examples/ has runnable recipes (they use ../../src directly, no install):
compare-two-commands,
find-a-hotspot,
heap-usage,
gate-a-regression,
profile-in-process,
benchmark-a-function.
cd examples/find-a-hotspot && bun run demo
bun run examples # all of them, from the repo root