Skip to content

fix(observability): record every documented metric, including notifications - #33

Merged
arg1998 merged 6 commits into
mainfrom
fix/otel-unrecorded-metrics
Sep 29, 2026
Merged

arg1998 merged 6 commits into
mainfrom
fix/otel-unrecorded-metrics

Conversation

@arg1998

@arg1998 arg1998 commented Sep 29, 2026

Copy link
Copy Markdown
Owner

The bug

A pre-release audit found metrics that are defined in the instrument catalogue and promised by spec 10 §7 and the telemetry guide, but never reach an OTLP collector. I re-checked it against main (7873a06):

  • Never recorded: browserhive.attention.wait, browserhive.session.launch.duration, browserhive.ws.connections, browserhive.ws.buffered_bytes, browserhive.ws.frames_dropped, browserhive.db.dropped_writes, browserhive.browser.rss_bytes and browserhive.process.event_loop_lag. The last one was also missing from the audit.
  • Documented attributes never set: closed_reason on session.lifetime, kind on attention.open, table on retention.pruned_rows.
  • The three notification counters were in fact recorded. build-domain passes them into the outbox, the act-button service and the report scheduler, and they showed up in the live check below. They were missing from the guide, though. notifications.reports also never emitted the documented manual outcome, and it emitted an undocumented revised for silent in-app anomaly revisions. Both are fixed. The composition now wires these counters through one helper, notificationCounters(), and the guard test uses the same helper.
  • Found during the live run: the SDK keeps sending the last value of an observable-gauge series after it stops being observed. So a closed session's browser memory, or a closed WebSocket's buffered bytes, would have been exported forever.
  • Before this change the consumers were wired even when --otelSignals left metrics out.

Spec first

The first commit makes spec 10 §7 the single source of truth: every instrument with its type, unit, attributes (closed value sets in parentheses) and when it is recorded. docs/guide/telemetry.md mirrors that table and now includes the notification and process metrics. Spec 09 names the guard test. Changes to the tables:

  • Units were added: ms, By, and UCUM annotations such as {call}.
  • process.* is split into three rows.
  • db.dropped_writes now lists the recorder tables, and every table is reported from the start at 0.
  • The manual outcome is defined for reports.
  • Browser memory is defined as the sum of the tree's RSS, on Linux and macOS.
  • The spec states that observable gauges report only the series observed at each export.

What was wired where

Metric Source
session.launch.duration {channel, stealth} first session.updated that carries launchMs (bus)
session.lifetime {closed_reason} session.closed (bus)
attention.open {kind} attention.* and vault.confirm.* (bus). Only requests seen opening are settled, so the startup reconcile can't push it below zero.
attention.wait {status} attention.resolved, from waited_ms
retention.pruned_rows {table} retention.completed, through a new internal prunedByTable key that the WS schema strips
ws.connections, ws.buffered_bytes (top 50), ws.frames_dropped {screencast, logs, feed} realtime hub, via realtimeMetrics(hub). The hub keeps plain dropped-frame totals.
db.dropped_writes {table} the write queue's new droppedWritesByTable
browser.rss_bytes {session_id} a 10 s sampler. It reads each session's browser pid once, via SessionHandle.browserPid() (a browser-level CDP SystemInfo.getProcessInfo; no page target is touched), then sums the process tree from /proc (Linux) or ps (macOS). Windows gets no data points.
process.event_loop_lag perf_hooks.monitorEventLoopDelay, reporting the p99 since the previous export

Observable gauges are now collected with delta temporality. OTLP gauges don't carry temporality, so the only effect is that series nobody observes anymore disappear. Counters stay cumulative. The app layer imports no infra, and dependency-cruiser is clean.

With OTel off, or with the metrics signal off, nothing is subscribed, sampled or timed. The only additions that always run are plain integer and map increments on the drop paths of the hub and the write queue.

Guard test

packages/browserhive/test/composition/metrics-guard.test.ts works in two steps.

  1. It parses both Markdown tables and requires spec == docs == METRIC_DEFINITIONS, row by row.

  2. A SOURCES table has to cover every catalogue metric exactly once. Each source is driven for real:

    • the bus consumers;
    • the SQLite write queue;
    • a real realtime hub with a socket;
    • the sampler reading a real process tree;
    • the event-loop monitor;
    • buildOps with a real webhook channel, an act-button press and an on-demand digest.

    Everything is exported through the real createTelemetry to an OTLP/HTTP receiver inside the test. Every metric must arrive with its documented kind, unit and attribute keys, and a closed session's gauge series must disappear.

I mutation-checked it. Each of these makes it fail: dropping a callback, dropping an attribute, adding an undocumented attribute, and reverting the delta-gauge change. There are also unit tests for every new piece, and an integration test that reads each real browser's pid and process-tree memory.

Real OTLP check

I ran the built daemon with --otel --otelProtocol=http/json against a small OTLP receiver, with an ntfy channel on FakePlatforms, a blocklist and the fake bw. The traffic was:

  • two MCP sessions;
  • a failing click;
  • a blocked navigate;
  • a vault_fill;
  • request_attention, answered by pressing the ntfy "Mark resolved" act button, plus a press of a token that was never issued;
  • a dashboard login and WebSocket;
  • an on-demand digest;
  • closing a session;
  • then SIGTERM.
22 metrics received in 4 exports:

browserhive.attention.open  [up-down counter, {request}]  1 pts  {"kind":"attention"}=0
browserhive.attention.wait  [histogram, ms]  1 pts  {"status":"resolved"}={"count":1,"sum":202}
browserhive.blocklist.hits  [counter, {hit}]  1 pts  {"source":"tool"}=1
browserhive.browser.rss_bytes  [gauge, By]  1 pts  {"session_id":"otela-2jutnqs5"}=1020932096
browserhive.db.dropped_writes  [counter, {write}]  8 pts  {"table":"sessions"}=0  {"table":"tool_calls"}=0  {"table":"pages"}=0
browserhive.db.size_bytes  [gauge, By]  1 pts  {}=446464
browserhive.db.write_queue.depth  [gauge, {write}]  1 pts  {}=0
browserhive.notifications.actions  [counter, {press}]  2 pts  {"channel_kind":"ntfy","outcome":"unknown"}=1  {"channel_kind":"ntfy","outcome":"done"}=1
browserhive.notifications.deliveries  [counter, {delivery}]  1 pts  {"channel_kind":"ntfy","status":"sent"}=4
browserhive.notifications.reports  [counter, {report}]  2 pts  {"kind":"digest.daily","outcome":"manual"}=1  {"kind":"digest.daily","outcome":"in_app"}=1
browserhive.process.event_loop_lag  [gauge, ms]  1 pts  {}=3.577855
browserhive.process.heap_bytes  [gauge, By]  1 pts  {}=71143067
browserhive.process.rss_bytes  [gauge, By]  1 pts  {}=229859328
browserhive.session.launch.duration  [histogram, ms]  1 pts  {"channel":"chromium","stealth":true}={"count":2,"sum":1162}
browserhive.session.lifetime  [histogram, ms]  1 pts  {"closed_reason":"user"}={"count":1,"sum":30786}
browserhive.sessions.active  [up-down counter, {session}]  1 pts  {"harness":"other"}=1
browserhive.tool_call.duration  [histogram, ms]  6 pts  {"tool":"launch_session"}={"count":2,"sum":1180}  {"tool":"navigate"}={"count":2,"sum":56}  {"tool":"click"}={"count":1,"sum":30007}
browserhive.tool_calls  [counter, {call}]  7 pts  {"tool":"launch_session","ok":true,"harness":"other"}=2  {"tool":"navigate","ok":true,"harness":"other"}=1  {"tool":"click","ok":false,"harness":"other","error_code":"ELEMENT_NOT_ACTIONABLE"}=1
browserhive.vault.fills  [counter, {fill}]  1 pts  {"result":"blocked"}=1
browserhive.ws.buffered_bytes  [gauge, By]  1 pts  {"connection_id":"c-DTjILvV3nD"}=0
browserhive.ws.connections  [up-down counter, {connection}]  1 pts  {}=1
browserhive.ws.frames_dropped  [counter, {frame}]  3 pts  {"channel":"screencast"}=0  {"channel":"logs"}=0  {"channel":"feed"}=0

22 of the 23 metrics arrived live. retention.pruned_rows can't show up in a short run, because retention first runs 6 h after start. The guard test covers it. A ps check of each browser tree (~1.06–1.09 GB) agrees with the metric, and the closed session no longer appeared in the next export.

Local gate

  • bun run check, test:goldens, build, package:check and license:check all pass.
  • test:integration: 68 passed, 4 skipped.
  • e2e, replicated like CI: 15 passed, 1 skipped.
  • The website builds.

Spec 10 §7 and the telemetry guide now list every instrument with its
type, unit, attributes and when it is recorded, including the three
notification counters and the three process gauges, and name the guard
test that keeps the tables, the catalogue and the recorded metrics in
step (spec 09 §3.2).
The adapter only fed eleven of the catalogue's instruments. It now
records the rest from their sources:

- session launch duration (channel, stealth) from the first
  session.updated carrying launch_ms; lifetime by closed_reason
- attention.open by kind (attention and vault confirmations) and the
  attention wait by status, from the broker's events
- retention pruned rows by table, carried internally on
  retention.completed
- WebSocket connections, buffered bytes (top 50) and frames dropped by
  channel, read from the realtime hub's running totals at export
- dropped writes by table, read from the write queue's totals
- each live session's browser process-tree RSS, sampled every 10 s from
  the browser pid (one browser-level DevTools read, cached) and /proc or
  ps
- the p99 event-loop delay since the previous export

Every instrument now carries the unit of its catalogue row. The report
scheduler counts on-demand digests as manual and no longer counts a
silent revision of an in-app anomaly alert. With telemetry off nothing
is subscribed, sampled or timed.
A table-driven guard parses the metric tables of spec 10 §7 and the
telemetry guide and requires them to match the instrument catalogue row
for row, then drives every source the composition root wires (bus
consumers, the SQLite write queue, the realtime hub, the browser-memory
sampler over a real process tree, the event-loop monitor, and the
notification outbox, act buttons and report scheduler built by
buildOps) and requires each instrument to reach a real OTLP/HTTP
receiver with its documented type, unit and attribute keys.

Unit tests cover the process-tree reader, the event-loop monitor, the
sampler, the hub's and the write queue's totals, the catalogue units and
the report outcomes; an integration test reads each real browser's pid
and process-tree memory.
…corder table

The SDK repeats the last value of an observable gauge's series that is
no longer observed, so a closed session's browser memory or a closed
WebSocket's buffered bytes was exported forever. Observable gauges are
now collected with delta temporality, which OTLP gauges do not carry, so
each export holds only what exists; counters stay cumulative.

browserhive.db.dropped_writes now reports every recorder table from the
start at 0, so a rate over the series works before the first drop.

Spec 10 §7 and the guide say both, and that browser memory is the sum
of the processes' RSS. The guard test checks that a closed session's
series disappears and that the spec's recorder tables match.
…orted

With --otelSignals traces,logs the meter is a no-op, so the bus
consumers, the browser-memory sampler and the event-loop monitor are no
longer started for nothing.
@arg1998
arg1998 marked this pull request as ready for review September 29, 2026 17:02
@arg1998
arg1998 merged commit f4ebe1a into main Sep 29, 2026
23 of 25 checks passed
@arg1998
arg1998 deleted the fix/otel-unrecorded-metrics branch September 29, 2026 17:58
arg1998 pushed a commit that referenced this pull request Oct 2, 2026
This PR was opened by the [Changesets
release](https://github.com/changesets/action) GitHub action. When
you're ready to do a release, you can merge this and the packages will
be published to npm automatically. If you're not ready to do a release
yet, that's fine, whenever you add more changesets to main, this PR will
be updated.


# Releases
## browserhive@0.2.0

### Minor Changes

- [#17](#17)
[`a4b5b8a`](a4b5b8a)
Thanks [@arg1998](https://github.com/arg1998)! - Choose which browser
sessions use, and run it inside Chromium's sandbox.

- **Sessions run inside Chromium's sandbox wherever your machine allows
it.** The new `--sandbox` setting (`BROWSERHIVE_SANDBOX`) defaults to
`auto`: each browser is tried with the sandbox on its first launch and
keeps it where it works (macOS, Windows, most Linux, and Google Chrome
on Ubuntu). Where it cannot (Ubuntu 23.10+ with the bundled browser,
running as root, Docker), sessions run as before, with one warning in
the log, and `browserhive doctor` explains why with the fix for your
machine. `--sandbox off` is exactly the old behaviour.
- **`--sandbox on` makes the sandbox a guarantee.** The server checks
the configured browser before it opens its port and refuses to start
(exit code 3) when it cannot sandbox, printing the reason and what to do
on your machine, easiest first: an installed browser that does sandbox,
an AppArmor profile, running as a normal user, or `--sandbox auto`.
- **A sandbox the machine cannot provide is now a clear error.**
`launch_session` with `launch_options: { chromiumSandbox: true }` on
such a host used to answer `INTERNAL_ERROR` with "retry with backoff";
it now answers `SANDBOX_UNAVAILABLE`, "retrying will not help", with the
channels that do work.
- **`browserhive init` shows the browsers on your machine and lets you
pick one.** It lists the bundled Chromium, an installed Google Chrome or
Microsoft Edge, whether each can run sandboxed, and the pros and cons of
each for your system. Enter keeps the current choice; a new choice is
saved to your config file after you confirm. It can also install Google
Chrome for you (Google's installer, administrator rights). In scripts:
`--channel chrome --yes`, `--installChrome`; nothing is asked without a
terminal.
- **`browserhive doctor` checks the browser you configured.** A
`defaultChannel=chrome` without Chrome installed is now a failure
instead of a green check. New checks: installed Chrome and Edge, a
browser that has moved more than one major version ahead of the tested
build, managed policies that block automation, the sandbox per browser
with the fix for your OS, and running as root or in a container. `doctor
--printApparmorProfile` prints (never installs) the profile that lets
the bundled browser sandbox on Ubuntu.
- **The dashboard shows browsers and the sandbox.** The System page
lists the browsers found, their versions and whether each runs
sandboxed; a session's Details tab shows its browser version and sandbox
state.
- A missing Google Chrome now names the command that installs it
(`browserhive init --installChrome`), and `--no-sandbox` suggests
`--sandbox`.

- [#20](#20)
[`15871c2`](15871c2)
Thanks [@arg1998](https://github.com/arg1998)! - Read environment
variables in `browserhive.config.json`, and see which variable every
value came from.

- **Keep secrets out of the config file.** A string in
`browserhive.config.json` can now contain `{env:NAME}`, and BrowserHive
reads that environment variable when it starts: `"authTokens":
"ci-runner:{env:CI_TOKEN}"`, `"otelHeaders": { "Authorization": "Bearer
{env:OTLP_TOKEN}" }`, `"otelEndpoint": "http://{env:OTLP_HOST}:4318"`.
The file can be committed and shared while tokens and per-machine values
stay in the environment. It works for every key, in whole values, inside
longer strings, in lists and in header maps; `"maxSessions":
"{env:MAX_SESSIONS}"` accepts exactly what `BROWSERHIVE_MAX_SESSIONS`
would.
- **Defaults for one file everywhere.** `{env:NAME:-default}` uses
`default` when `NAME` is not set or is empty, so the same file works on
a laptop and on a server. `{env:NAME}` without a default is required: if
the variable is missing, startup stops and says which key, which file
and which variable, and how to add a default.
- **Every place shows where a value came from.** The startup log reads
`config: otelEndpoint=… (config-file via $OTLP_HOST) shadows env=…`,
`browserhive config show` prints `config-file via $OTLP_HOST` in the
SOURCE column, `config show --json` and `GET /api/v1/system/config` add
`refs` and `template` fields, and the dashboard's System page shows a
`$OTLP_HOST` chip next to the source. Select it to see the value as
written in the file; a new **Only values from references** switch lists
just those keys.
- **Secrets stay hidden.** For `authTokens`, `otelHeaders`, and any
value whose variable name looks like a credential (it contains `token`,
`secret`, `password` and the like), you see the variable's name, never
its value, and the value is scrubbed from logs. `browserhive doctor` no
longer asks you to `chmod 600` a config file whose `authTokens` only
reference variables, and it warns when a variable was not set so its
default is in use.
- **Mistakes are caught, not ignored.** `${env:NAME}` (the OpenTelemetry
Collector's spelling) stops startup with "Did you mean '{env:NAME}'?";
`{ENV:NAME}` or `{file:…}` stop it too, with how to keep such text
literal (`{{…}}`). Braces that are not references, like `{trace_id}` in
`otelTraceUrlTemplate`, are left alone, so existing config files work as
before. References are not expanded in `BROWSERHIVE_*` variables or
command-line flags (your shell does that); BrowserHive warns if it finds
one there.
- The JSON Schema for the config file (`browserhive config schema`)
accepts a reference for every key, so editors no longer underline
`"stealth": "{env:STEALTH}"`.

- [#21](#21)
[`43d2f5e`](43d2f5e)
Thanks [@arg1998](https://github.com/arg1998)! - See which agent is
connected (Claude Code, Codex, Cursor, OpenCode, Gemini CLI and others),
count sessions and tool calls per agent, and filter by it.

- **Recognised with nothing to configure.** Claude Code and Gemini CLI
over stdio, and Claude Code, Codex, OpenCode, Cursor, VS Code, Cline,
Continue and Zed over HTTP, are recognised from what they already send.
Anything BrowserHive can't place is shown as **Unknown**, which is
counted and filterable like any other agent, never an empty cell.
- **Name any agent in one line.** Add an `X-BH-Agent-Harness` header, a
`BROWSERHIVE_HARNESS` variable in a stdio server's `env` block, or
`?harness=<name>` to the MCP URL when a client only has a URL field.
`X-BH-Agent-Model` / `BROWSERHIVE_MODEL` and `X-BH-Workspace` /
`BROWSERHIVE_WORKSPACE` label the model and the workspace;
`X-BH-Meta-<Name>` headers add a few extra labels. Tool calls can carry
the same in `_meta` (`ai.browserhive/harness`, `ai.browserhive/model`,
`ai.browserhive/workspace`), read on every call.
- **In the dashboard.** The sessions list has a Harness filter and an
Agent column; a session's Details tab has a Client panel (the agent and
how it was recognised, the model or "not reported", the workspace, the
client's name, version and protocol); the Overview has a Harnesses card
with sessions and tool calls per agent; the System page lists live and
recent MCP connections, with the User-Agent, IP, conflicting signals and
extra labels of each. A tool call's details show which agent made it.
- **In the API and telemetry.** `GET /api/v1/sessions` accepts
`harness=` and returns a `harnesses` facet, each session has a
`harness`, tool calls carry `harness` (and `GET /api/v1/tool-calls`
filters by it), and there are two new endpoints: `GET
/api/v1/metrics/harnesses` and `GET /api/v1/system/mcp/connections`.
Traces carry `browserhive.harness` (and the declared model); the
tool-call and live-session metrics gain a `harness` attribute limited to
known names plus `other` and `unknown`.
- **Reported, not verified.** All of this is what the client or your own
configuration says. BrowserHive shows it faithfully and never uses it to
allow or refuse anything. MCP gives a server no way to learn the model,
so a model is shown only when one is declared.
- Existing sessions and databases keep working: the database upgrades on
start (schema v3), sessions from before read Unknown, and every field
that was there before is still there. `BROWSERHIVE_HARNESS`,
`BROWSERHIVE_MODEL` and `BROWSERHIVE_WORKSPACE` are not configuration
settings, so they are no longer rejected as unknown, and a misspelled
one gets a suggestion.

- [#28](#28)
[`9470feb`](9470feb)
Thanks [@arg1998](https://github.com/arg1998)! - Answer from your phone:
Approve, Reject and Mark resolved right in Telegram, Discord and ntfy,
and richer Telegram messages.

- **Answer without opening the dashboard.** Switch on **Answer from the
chat** for a channel, and a notification that waits for you carries
buttons that act: **Mark resolved** and **Reject** for an attention
request, **Approve** and **Deny** for a vault fill. Press one and
BrowserHive does what the same button in the dashboard does, then edits
the message: the buttons disappear and it says who answered ("Resolved
on Telegram by … after 42 s"). Off by default.
- **Only the right person, only once.** On Telegram and Discord only the
accounts on the channel's allow-list may press; by default that is the
person who connected the chat, and anyone else is told their id so you
can add them. Every button works once, for 24 hours, only in its own
chat, and only while the request still waits. Every press is listed
under the new **Notifications → Actions**. The agent learns that you
answered from Telegram, Discord or ntfy, never your chat identity.
- **Discord bot mode.** A Discord channel can now use a bot instead of a
webhook, the only way to press buttons in Discord. The wizard walks
through the Developer Portal, builds the invite link with the minimal
permissions, lists your servers and channels, and links your account
with a **This is me** button. The channel card shows whether the bot is
connected.
- **ntfy answers through a second topic.** Give an ntfy channel a reply
topic, and its buttons make your phone post the answer there;
BrowserHive listens and updates the notification. Works on Android and
iOS.
- **Nothing to expose.** Every connection goes out from your machine
(Telegram long polling, the Discord gateway, an ntfy subscription).
Presses made while BrowserHive was stopped are handled at the next start
when the platform kept them. The webhook channel carries the act actions
as they are, and your receiver answers through the REST API.
- **Telegram Rich Messages.** Telegram notifications now have a heading,
the facts as a table, real tables, collapsible quotes and coloured
buttons, with the screenshot inside the message. If Telegram refuses
one, the classic format is sent instead.
- **Setup and terminal.** Startup channels take `actButtons=true` and
`allow=<user ids>`, `mode=bot` for Discord and `reply=` for ntfy.
`browserhive channels list` shows whether answers reach BrowserHive.
- New REST endpoints: `GET /api/v1/channels/actions` and the Discord bot
setup under `/api/v1/channels/discord/…`; channels report `connection`;
the `channels` WebSocket topic adds `action.recorded`. The database
moves to schema v6 (three new tables; an older release still opens it).

- [#27](#27)
[`fb94fbc`](fb94fbc)
Thanks [@arg1998](https://github.com/arg1998)! - Notifications on your
phone: Telegram, Discord, ntfy and webhook channels, with screenshots,
live updates and messages that delete themselves.

- **Get a message when an agent needs you.** Add a channel under
**Notifications → Channels**: a Telegram bot (one-tap connect, no chat
id to look up), a Discord webhook, an ntfy topic (scan a QR code with
the ntfy app) or a webhook of your own. A wizard shows the exact line to
set the token for how you run BrowserHive, checks that it is set,
previews the message exactly as it will look, and sends a test. Presets
pick what to send (_Needs me now_, _Problems_, _Wrap-ups_); **Advanced**
adds minimum severity, session patterns, harness, quiet hours with a
time zone and the content level.
- **Messages keep up.** When you resolve an attention request, the chat
message is edited in place, silently, and its buttons disappear; a
growing group of tool errors updates its count. Every send, edit and
delete is in the new **Delivery log**, live, with a sentence for
anything that was not sent ("quiet hours", "the platform refused the
token").
- **Screenshots, when you want them.** Off by default, per channel and
category: the page when an agent asked for help (CAPTCHAs included), the
login page before a vault fill (never during one), a crashed session's
last frame. Form fields can be masked.
- **Self-destruct.** Delete messages after a time you choose per
category, or once they are resolved. Telegram only allows 48 hours, so
its timers stop at 47.
- **Links that open on your phone.** The new `publicUrl` key
(`--publicUrl`, `BROWSERHIVE_PUBLIC_URL`) is the address where you reach
the dashboard (a Tailscale name, your reverse proxy, a Cloudflare
tunnel). Notification links use it, its host is trusted without
`allowedHosts`, and the CSRF check accepts it even when your proxy
rewrites `Host`. The System page and `browserhive doctor` check that it
really reaches this BrowserHive.
- **Your accounts, your tokens.** BrowserHive runs no servers or shared
bots. Tokens stay in environment variables; channels store only the
variable names, so a database backup never contains one.
- **For servers and containers**, declare channels at startup with
`--notificationChannel
"telegram:name=phone,token=env:BH_TG_TOKEN,chat=123456"` (repeatable; a
token typed into the flag is refused). They show in the dashboard with a
"from startup" badge.
- **From a terminal:** `browserhive channels list`, `channels test
<name>` and `channels preview <name>`; `browserhive doctor` checks every
channel's variables and `publicUrl`.
- New REST endpoints under `/api/v1/channels` (scopes `channels:read`,
`channels:write`), `GET /api/v1/system/public-url`, a `channels`
WebSocket topic, and `instance_id` in `GET /health`.

- [#25](#25)
[`9fce259`](9fce259)
Thanks [@arg1998](https://github.com/arg1998)! - Notifications now
follow what they announce, and every notification is a versioned message
that can be delivered reliably to other apps.

- **A notification keeps its place as things change.** When you resolve
an attention request or a vault confirmation, reject it, or it times
out, its notification shows the outcome (**resolved**, **expired**) as a
small pill in the bell and on the Notifications page, instead of staying
as if it were still waiting. A toast still on screen for it closes. A
recovered subsystem marks its degradation notification resolved the same
way.
- **Richer notification data in the API.** `GET /api/v1/notifications`
and the `notifications` WebSocket topic add `kind`, `category`,
`severity`, `state`, `revision` and `thread` to every notification.
Existing fields are unchanged.
- **Safer text.** An agent's attention reason and other text copied into
a notification now go through the same redaction as the logs, and page
addresses lose their query strings.
- **Built for delivery to your phone.** Every notification is also a
versioned message document, whose JSON Schema is published in the
reference docs (the webhook channel sends it as is). Delivery to
Telegram, Discord, ntfy and webhooks goes through a new outbox in the
database, with retries and a circuit breaker, so nothing is lost when a
service is down. With no channel configured nothing extra runs.
- The database upgrades on start (schema v5: new columns on
notifications and three new tables; a backup is written first). Older
releases can still open it. Existing notifications are classified from
what they already recorded; nothing is invented for them.

- [#29](#29)
[`7873a06`](7873a06)
Thanks [@arg1998](https://github.com/arg1998)! - A daily summary and a
heads-up when something's off.

- **Daily or weekly digest.** Give a channel a digest (every day at
09:00, optionally weekdays only with Monday covering the weekend, or
every week on Friday at 17:00; day and time are yours to change) and it
gets the period in numbers: sessions, tool calls and errors with the
rate, attention requests and how fast they were answered, vault fills,
blocked requests, the slowest tool against the period before, the top
errors, open problems, a small chart of tool calls per hour and a table
per harness, with a link to the Overview for exactly that period. The
new **Daily digest** preset sets it up in one click.
- **In your time zone.** Each channel has a time zone, BrowserHive's own
unless you pick another, and digests and quiet hours follow it through
daylight saving time.
- **Nothing lost, nothing spammed.** If BrowserHive was off when a
digest was due, the most recent one arrives when it starts again, marked
late, with how many earlier ones were skipped. A day with no activity
sends nothing (the delivery log says so). A digest due in quiet hours
arrives silently.
- **Tell me when something looks off.** An hourly check that stays
silent until a threshold is crossed: many tool calls failing, a request
waiting too long, sessions at the limit, a spike in blocked requests,
BrowserHive degraded. The alert updates itself and says **Back to
normal** when things recover, without flapping. Thresholds can be tuned
per channel.
- **Reports in BrowserHive.** Every digest and anomaly alert also lands
in the dashboard once per period, however many channels it reached: in
the bell and the inbox (a new **Reports** filter), quietly for digests
(no pop-up, no badge) and like a System notification for anomaly alerts.
The new **Notifications → Reports** tab keeps them for 90 days, even
after you dismiss them, with filters and a page per report (its numbers,
chart and tables, the channels it reached, and Open Overview for this
period). It can also run a digest and anomaly alerts for the dashboard
alone, with no channel at all.
- **Send a digest now.** Preview the real digest exactly as your phone
will show it, then send it on demand, from the channel card or `POST
/api/v1/channels/{id}/digest`.
- **Setup and terminal.** Startup channels take `digest=daily@09:00`,
`digest=daily:weekdays` or `digest=weekly` (Friday 17:00;
`weekly:mon@08:30` for another day), `tz=` and `anomaly=on` with
`anomaly.*` thresholds; `browserhive channels list` shows each channel's
next digest, and `channels preview --sample digest|anomaly` renders the
samples.
- **Also:** the Allow this person button explains when you lack
`channels:write`, a deleted channel's reply-topic cursor goes with it,
the live delivery log no longer shows stale superseded rows, and the
public address check reports a proxy's 5xx page as unreachable. The
message contract gains the `digest.weekly` kind, a `chart` block and an
optional `report` field (schema 1, additive); `GET
/api/v1/notifications` takes a `category` filter; new `GET
/api/v1/notifications/reports`, `GET /api/v1/notifications/reports/{id}`
and `GET`/`PUT /api/v1/notifications/report-settings`; no database
migration.

- [#24](#24)
[`6c90ade`](6c90ade)
Thanks [@arg1998](https://github.com/arg1998)! - Closed sessions now
keep showing whether they ran inside Chromium's sandbox, and with which
browser version.

- **Recorded at launch.** When a session's browser starts, BrowserHive
stores whether it runs sandboxed and the browser's real version with the
session. Under the default `--sandbox auto` the answer depends on the
browser and the machine (on Ubuntu, Google Chrome sandboxes and the
bundled Chromium falls back), so it is recorded rather than worked out
later.
- **In the dashboard.** A session's Details tab shows the browser
version and a **sandboxed** / **not sandboxed** state for finished
sessions too, not only while they run. Sessions from before this release
show **not recorded**, with a note that this does not mean the sandbox
was off; a session whose browser never started shows **not launched**.
- **In the API.** `browser: { version, sandboxed }` on `GET
/api/v1/sessions` and `GET /api/v1/sessions/{id}` is now filled for
closed sessions from what was recorded at launch. It is still left out
when nothing was recorded, so read a missing `browser` as "unknown",
never as "not sandboxed".
- The database upgrades on start (schema v4, two new columns, a backup
is written first); older releases can still open it. Nothing is guessed
for existing sessions.

### Patch Changes

- [#22](#22)
[`a7a79d8`](a7a79d8)
Thanks [@arg1998](https://github.com/arg1998)! - The System page's **MCP
connections** list now shows 10 connections per page, with the usual
pager underneath (10, 25 or 50 rows per page, previous/next, "1–10 of
23"). Before, it showed up to 50 at once with no way to see older ones.
For API clients, `GET /api/v1/system/mcp/connections` accepts an
`offset` for paging and returns `total`, the number of stored
connections; existing calls behave exactly as before.

- [#33](#33)
[`f4ebe1a`](f4ebe1a)
Thanks [@arg1998](https://github.com/arg1998)! - OpenTelemetry now
exports the notification, attention, session-launch, WebSocket,
write-queue and browser-memory metrics the docs describe.

- **Metrics that were documented but never sent** now reach your
collector: `browserhive.attention.wait`,
`browserhive.session.launch.duration`, `browserhive.ws.connections`,
`browserhive.ws.buffered_bytes`, `browserhive.ws.frames_dropped`,
`browserhive.db.dropped_writes`, `browserhive.browser.rss_bytes` (each
session's browser with all its processes, every 10 seconds; Linux and
macOS) and `browserhive.process.event_loop_lag`.
- **Attributes the docs promised** are now set: `closed_reason` on
`browserhive.session.lifetime`, `kind` on `browserhive.attention.open`
(vault confirmations count too), `table` on
`browserhive.retention.pruned_rows`.
- **The notification metrics** `browserhive.notifications.deliveries`,
`.actions` and `.reports` are now in the [telemetry
guide](https://browserhive.ai/docs/guide/telemetry). `.reports` now
counts on-demand digests as `manual`, as documented, and no longer
counts a silent revision of an open in-app anomaly alert.
- **Gauges only report what exists**: a closed session's browser memory
or a closed WebSocket's buffered bytes disappears from the next export
instead of repeating its last value. `browserhive.db.dropped_writes`
reports every recorder table from the start, at 0.
- **Every metric has a unit** (`ms`, `By`, or a count such as `{call}`),
and the guide's table lists each one with its type, unit and attributes.
With telemetry off nothing is measured, as before.

- [#15](#15)
[`280ec1f`](280ec1f)
Thanks [@arg1998](https://github.com/arg1998)! - Fixes found while
researching the next features.

- **MCP works behind a port mapping or SSH tunnel.** `/mcp` rejected
every request whose `Host` named a different port (`localhost:8080` for
a server on 9876) while the dashboard kept working. Both now use the
same check, which ignores the port.
- **New `--allowedHosts` setting** (`BROWSERHIVE_ALLOWED_HOSTS`) for the
name a reverse proxy forwards, so the recommended TLS proxy setup works
without rewriting `Host`.
- **Session details show the MCP client that launched the session** —
its name and version, plus the model and workspace when it sends
`X-BH-Agent-Model` / `X-BH-Workspace`. `X-BH-Agent-Harness` no longer
overwrites the workspace. Stdio connections record their client too.
- **Saved storage states keep IndexedDB**, so sites that store their
login there (Firebase Auth, among others) restore logged in.
- **A malformed secret setting is no longer printed in the error.**
`BROWSERHIVE_AUTH_TOKENS` and `otelHeaders` values from env or the
config file are also scrubbed from logs and telemetry now.
- **Takeover input is audited**: one audit row per session and second
with counts only, never the keys typed.
- **`maxSessions` respects container and systemd memory limits** instead
of deriving from the host's full RAM.
- **A data directory on a filesystem that refuses `chmod`** (network
shares, some bind mounts) no longer stops the server from starting; it
logs a warning.
- **Toggling fullscreen in the live view no longer restarts the
stream.**
  - Old MCP connection records are now pruned by retention.
- A helper process (such as the Bitwarden CLI) that exits before reading
its input no longer raises a spurious `UNHANDLED` degradation.
## @browserhive/core@0.2.0

### Patch Changes

- Updated dependencies []:
  - @browserhive/contracts@0.2.0
## @browserhive/dashboard@0.2.0

### Patch Changes

- Updated dependencies []:
  - @browserhive/contracts@0.2.0
## @browserhive/contracts@0.2.0

No changes in this release.

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant