Skip to content

Release/0.7.6 - #35

Merged
maltsev-dev merged 3 commits into
masterfrom
release/0.7.6
Jun 27, 2026
Merged

Release/0.7.6#35
maltsev-dev merged 3 commits into
masterfrom
release/0.7.6

Conversation

@maltsev-dev

Copy link
Copy Markdown
Member

What

Why

How

Test plan

  • Unit tests pass (per-repo, e.g. cd backend && cargo test, cd frontend && npm test)
  • Lint passes (per-repo, e.g. cd frontend && npm run lint)
  • Type-check passes (per-repo, e.g. cd frontend && npm run type-check)
  • Manually verified in dev / staging

Risk

Checklist

  • I have read the repo's CONTRIBUTING.md (if present)
  • My change does not introduce new lint warnings
  • I have updated the CHANGELOG (if user-visible)
  • I have considered backwards compatibility

Additive patch on top of the 0.7.0 thin-client refactor. No
breaking changes.

Added
-----

* nullrun.integrations.fastapi — one-line FastAPI integration
  that turns every NullRunDecision / NullRunInfrastructureError
  thrown by @nullrun.protect endpoints into a clean JSON
  response with the right HTTP status code. No per-endpoint
  except blocks required.

  Response shape:
    {"error_code": "NR-B004",
     "user_message": "You've reached the usage limit...",
     "category": "decision"}

  HTTP status mapping:
    * NR-B004 (budget), NR-L001 (loop), NR-R001 (rate) -> 429
      with optional Retry-After
    * NR-T001 (tool blocked), NR-X001 (generic block) -> 403
    * NR-W003 (paused) -> 503 with Retry-After
    * NR-W002 (killed) -> 503; WorkflowKilledInterrupt is a
      BaseException subclass so Starlette's
      add_exception_handler refuses it — handled via ASGI
      middleware instead (hybrid pattern, documented in
      module docstring).
    * NullRunInfrastructureError subclasses -> 503 (our side,
      not user's).

* nullrun.messages — default user-facing message catalog.
  Every NR-* error code has an English default message owned
  by NULLRUN, not customer code. Customer Support Bots hitting
  a budget cap show the same wording across every NullRun-backed
  application.
    * format_user_message(exc) — render exception as user-facing
      string
    * set_user_message(code, text) — per-process override for
      branded variants
    * get_user_message(code) — raw lookup
    * reset_overrides() — clear all overrides (for tests)

Changed
-------

* Transport._send_batch canonical JSON serialization — route the
  /track/batch body through _signed_request_body for consistent
  compact-separator serialization. HMAC itself is unaffected,
  but consistent serialization removes a special-case from the
  wire-format contract tests.

* Transport._send_batch actions response handling — backend
  renamed BatchTrackResponse.actions_taken (debug names) ->
  BatchTrackResponse.actions (ActionTaken structs). Read both
  for forward-compat; per-element try/except so one malformed
  entry doesn't abort the whole loop.

* pyproject.toml metadata — long-form description with search
  keywords, Maintainer: populated via maintainers=[...],
  expanded classifiers (Linux / Windows / macOS, Python 3.13,
  CPython, Security / AI / WWW/HTTP topics), project URL
  expander.

Tests
-----

* tests/test_messages.py (new, 282 lines) — catalog
  completeness (every NR-* code has a default message),
  override / reset behavior, render path.
* tests/test_integrations_fastapi.py (new, 289 lines) — HTTP
  status mapping per error code, response shape, ASGI
  middleware path for WorkflowKilledInterrupt, hybrid
  composition.
* tests/test_decision_split.py (new, 199 lines) — pins the
  decision / infrastructure error split.
* Updates to tests/test_runtime.py, tests/test_extractors.py
  reflecting transport canonical-JSON + actions-renamed
  changes.

Release plumbing
----------------

* pyproject.toml: version bumped 0.7.0 -> 0.7.6
* src/nullrun/__version__.py: __version__ = "0.7.6"
* CHANGELOG.md: full 0.7.6 entry covering additions,
  transport changes, metadata improvements

Tests pass locally (per session log) — pytest on Windows /
Python 3.14.2 is green.
@codecov

codecov Bot commented Jun 27, 2026

Copy link
Copy Markdown

…padding

PR #35 (release/0.7.6) failed all four CI jobs (test 3.10/3.11/3.12,
coverage, codecov/patch) on the same root cause + one latent bug
masked by it. This commit lands the fixes plus the last-mile tests
that bring coverage above the 82% threshold.

CI failure root
---------------

* tests/test_integrations_fastapi.py does from fastapi import ...
  at module top-level. CI installs only pip install -e '.[dev]',
  and fastapi was declared as an *optional* [fastapi] extra,
  NOT in [dev]. Pytest collection aborted with
  ModuleNotFoundError: No module named 'fastapi' → all 4 jobs red.
* Fix: add fastapi>=0.100,<1.0 to [dev]. Same precedent as
  langchain-core (already in [dev] for the same import-time
  contract: nullrun.instrumentation.langgraph is eager-imported
  from nullrun.decorators at collection time, so the test extras
  must cover the import chain).

Latent bug surfaced by the first fix
------------------------------------

The same PR refactored Transport._send_batch_with_retry_info to
route the /track/batch body through _signed_request_body for
canonical-JSON serialization (matching /gate and /execute). The two
sibling call sites use the module-level helper _signed_request_body
(no self.); this one used self._signed_request_body by typo.
Result: AttributeError on every batch flush, breaking 15 existing
tests across test_transport.py / test_track_batch_retry.py /
test_integration_contract.py / test_signal_safety.py. As long as
the fastapi collection error aborted pytest, this was hidden. Fixed
to _signed_request_body(...) with a docstring noting why it is
module-level and what the bug looked like.

Coverage padding (codecov/patch was failing on this too)
--------------------------------------------------------

Total coverage on the failing CI run was 81.98% — 0.02pp under the
fail-under=82 gate. After the two fixes above it would have
recovered to ~82.0% on the dot, so I added minimal tests for the
cheapest-to-cover gaps:

* tests/test_breaker_main.py (new) — covers the 5 statements in
  nullrun.breaker.__main__.main() (0% → 100%). The module
  exists so python -m nullrun.breaker exits cleanly instead of
  failing with No module named nullrun.breaker.__main__; the
  previous fix-mechanism was return 0 after a print, but no
  test was exercising it.
* tests/test_status.py — extends TestSummary with seven
  scenarios covering each conditional branch of NullRunStatus.summary()
  (organization_id, workflow_id, workflow_state != Normal,
  backend_reachable=False, ws_connected=False, recent_errors).
  status.py jumps 84.52% → 98.81%.
* tests/test_integrations_fastapi.py — four tests on
  _build_headers covering non-numeric, zero, negative, and
  resume_after (the WorkflowPausedException code path).
  integrations/fastapi.py jumps 90.22% → 94.57%.

After all three: TOTAL 81.98% → 82.46%, comfortably above the gate.

Verification
------------

* Local pytest: 997 passed, 13 skipped, 0 failed
  (Windows / Python 3.14.2, 8m47s — same env the original commit
  was validated in).
* python -m coverage report — 82.46%, no fail-under complaint.
…ng/tools

Patch coverage on PR #35 was 62.38% against a 65% threshold (codecov
target 70% / threshold 5pp). The two biggest delta-holders against
master were auto.py (+286) and langgraph.py (+221), both dominated
by Phase 4.1 additions:

  * auto._normalize_finish_reason + _FINISH_REASON_MAP
  * auto._openai_extractor  second-tier fields (cache_read_tokens,
    cache_write_tokens, reasoning_tokens, finish_reason, tool_names)
  * auto._anthropic_extractor cache_read / cache_write
  * langgraph._safe_get_gen_message
  * langgraph._get_finish_reason (5-source fallback chain)
  * langgraph.extract_usage_from_response second-tier fields

These are pure / near-pure functions with no network or vendor SDK
calls. Coverage padding is cheap — pin the canonical wire shapes
once and the backend ingest contract gets a free live spec.

Local numbers:
  * auto.py        63.44% -> 64.01%   (file-level, +57 statements)
  * langgraph.py   78.50% -> 86.01%   (file-level, +32 statements)
  * TOTAL          82.46% -> 83.13%   (already above 82% gate)

41 tests, all green. Existing test_extractors.py and
test_langgraph_callback.py left untouched — these tests
deliberately target the Phase 4.1 fields (cache_read /
cache_write / reasoning / finish_reason / tool_names) that the
older tests didn't pin.
@maltsev-dev
maltsev-dev merged commit 8fb8dbc into master Jun 27, 2026
4 checks passed
@maltsev-dev
maltsev-dev deleted the release/0.7.6 branch June 27, 2026 10:32
maltsev-dev added a commit that referenced this pull request Jun 27, 2026
Conflict resolution between release/0.7.7 (T4 per-call context for
/gate) and origin/master (Release/0.7.6 #35, which bumped the SDK
to 0.7.6):

* pyproject.toml: keep 0.7.7 (the HEAD side). 0.7.6 on master is
  superseded by 0.7.7 once this merges.
* CHANGELOG.md: keep BOTH the new 0.7.7 block (from HEAD) and the
  0.7.6 block (from master). They document different releases and
  are listed in chronological order with the older 0.7.6 block below.
* src/nullrun/{__init__.py, runtime.py, transport.py}: auto-merged
  cleanly - master doesn't touch the T4 hunks.

Auto-merge result equals HEAD, but the merge commit is still
needed to record the parent relationship and clear the conflict
state on the PR.
maltsev-dev added a commit that referenced this pull request Jun 27, 2026
…36)

* release: 0.7.6 — FastAPI integration + user-facing message catalog

Additive patch on top of the 0.7.0 thin-client refactor. No
breaking changes.

Added
-----

* nullrun.integrations.fastapi — one-line FastAPI integration
  that turns every NullRunDecision / NullRunInfrastructureError
  thrown by @nullrun.protect endpoints into a clean JSON
  response with the right HTTP status code. No per-endpoint
  except blocks required.

  Response shape:
    {"error_code": "NR-B004",
     "user_message": "You've reached the usage limit...",
     "category": "decision"}

  HTTP status mapping:
    * NR-B004 (budget), NR-L001 (loop), NR-R001 (rate) -> 429
      with optional Retry-After
    * NR-T001 (tool blocked), NR-X001 (generic block) -> 403
    * NR-W003 (paused) -> 503 with Retry-After
    * NR-W002 (killed) -> 503; WorkflowKilledInterrupt is a
      BaseException subclass so Starlette's
      add_exception_handler refuses it — handled via ASGI
      middleware instead (hybrid pattern, documented in
      module docstring).
    * NullRunInfrastructureError subclasses -> 503 (our side,
      not user's).

* nullrun.messages — default user-facing message catalog.
  Every NR-* error code has an English default message owned
  by NULLRUN, not customer code. Customer Support Bots hitting
  a budget cap show the same wording across every NullRun-backed
  application.
    * format_user_message(exc) — render exception as user-facing
      string
    * set_user_message(code, text) — per-process override for
      branded variants
    * get_user_message(code) — raw lookup
    * reset_overrides() — clear all overrides (for tests)

Changed
-------

* Transport._send_batch canonical JSON serialization — route the
  /track/batch body through _signed_request_body for consistent
  compact-separator serialization. HMAC itself is unaffected,
  but consistent serialization removes a special-case from the
  wire-format contract tests.

* Transport._send_batch actions response handling — backend
  renamed BatchTrackResponse.actions_taken (debug names) ->
  BatchTrackResponse.actions (ActionTaken structs). Read both
  for forward-compat; per-element try/except so one malformed
  entry doesn't abort the whole loop.

* pyproject.toml metadata — long-form description with search
  keywords, Maintainer: populated via maintainers=[...],
  expanded classifiers (Linux / Windows / macOS, Python 3.13,
  CPython, Security / AI / WWW/HTTP topics), project URL
  expander.

Tests
-----

* tests/test_messages.py (new, 282 lines) — catalog
  completeness (every NR-* code has a default message),
  override / reset behavior, render path.
* tests/test_integrations_fastapi.py (new, 289 lines) — HTTP
  status mapping per error code, response shape, ASGI
  middleware path for WorkflowKilledInterrupt, hybrid
  composition.
* tests/test_decision_split.py (new, 199 lines) — pins the
  decision / infrastructure error split.
* Updates to tests/test_runtime.py, tests/test_extractors.py
  reflecting transport canonical-JSON + actions-renamed
  changes.

Release plumbing
----------------

* pyproject.toml: version bumped 0.7.0 -> 0.7.6
* src/nullrun/__version__.py: __version__ = "0.7.6"
* CHANGELOG.md: full 0.7.6 entry covering additions,
  transport changes, metadata improvements

Tests pass locally (per session log) — pytest on Windows /
Python 3.14.2 is green.

* ci: fix PR #35 — fastapi dep + Transport._send_batch typo + coverage padding

PR #35 (release/0.7.6) failed all four CI jobs (test 3.10/3.11/3.12,
coverage, codecov/patch) on the same root cause + one latent bug
masked by it. This commit lands the fixes plus the last-mile tests
that bring coverage above the 82% threshold.

CI failure root
---------------

* tests/test_integrations_fastapi.py does from fastapi import ...
  at module top-level. CI installs only pip install -e '.[dev]',
  and fastapi was declared as an *optional* [fastapi] extra,
  NOT in [dev]. Pytest collection aborted with
  ModuleNotFoundError: No module named 'fastapi' → all 4 jobs red.
* Fix: add fastapi>=0.100,<1.0 to [dev]. Same precedent as
  langchain-core (already in [dev] for the same import-time
  contract: nullrun.instrumentation.langgraph is eager-imported
  from nullrun.decorators at collection time, so the test extras
  must cover the import chain).

Latent bug surfaced by the first fix
------------------------------------

The same PR refactored Transport._send_batch_with_retry_info to
route the /track/batch body through _signed_request_body for
canonical-JSON serialization (matching /gate and /execute). The two
sibling call sites use the module-level helper _signed_request_body
(no self.); this one used self._signed_request_body by typo.
Result: AttributeError on every batch flush, breaking 15 existing
tests across test_transport.py / test_track_batch_retry.py /
test_integration_contract.py / test_signal_safety.py. As long as
the fastapi collection error aborted pytest, this was hidden. Fixed
to _signed_request_body(...) with a docstring noting why it is
module-level and what the bug looked like.

Coverage padding (codecov/patch was failing on this too)
--------------------------------------------------------

Total coverage on the failing CI run was 81.98% — 0.02pp under the
fail-under=82 gate. After the two fixes above it would have
recovered to ~82.0% on the dot, so I added minimal tests for the
cheapest-to-cover gaps:

* tests/test_breaker_main.py (new) — covers the 5 statements in
  nullrun.breaker.__main__.main() (0% → 100%). The module
  exists so python -m nullrun.breaker exits cleanly instead of
  failing with No module named nullrun.breaker.__main__; the
  previous fix-mechanism was return 0 after a print, but no
  test was exercising it.
* tests/test_status.py — extends TestSummary with seven
  scenarios covering each conditional branch of NullRunStatus.summary()
  (organization_id, workflow_id, workflow_state != Normal,
  backend_reachable=False, ws_connected=False, recent_errors).
  status.py jumps 84.52% → 98.81%.
* tests/test_integrations_fastapi.py — four tests on
  _build_headers covering non-numeric, zero, negative, and
  resume_after (the WorkflowPausedException code path).
  integrations/fastapi.py jumps 90.22% → 94.57%.

After all three: TOTAL 81.98% → 82.46%, comfortably above the gate.

Verification
------------

* Local pytest: 997 passed, 13 skipped, 0 failed
  (Windows / Python 3.14.2, 8m47s — same env the original commit
  was validated in).
* python -m coverage report — 82.46%, no fail-under complaint.

* test: cover Phase 4.1 instrumentation — finish_reason + cache/reasoning/tools

Patch coverage on PR #35 was 62.38% against a 65% threshold (codecov
target 70% / threshold 5pp). The two biggest delta-holders against
master were auto.py (+286) and langgraph.py (+221), both dominated
by Phase 4.1 additions:

  * auto._normalize_finish_reason + _FINISH_REASON_MAP
  * auto._openai_extractor  second-tier fields (cache_read_tokens,
    cache_write_tokens, reasoning_tokens, finish_reason, tool_names)
  * auto._anthropic_extractor cache_read / cache_write
  * langgraph._safe_get_gen_message
  * langgraph._get_finish_reason (5-source fallback chain)
  * langgraph.extract_usage_from_response second-tier fields

These are pure / near-pure functions with no network or vendor SDK
calls. Coverage padding is cheap — pin the canonical wire shapes
once and the backend ingest contract gets a free live spec.

Local numbers:
  * auto.py        63.44% -> 64.01%   (file-level, +57 statements)
  * langgraph.py   78.50% -> 86.01%   (file-level, +32 statements)
  * TOTAL          82.46% -> 83.13%   (already above 82% gate)

41 tests, all green. Existing test_extractors.py and
test_langgraph_callback.py left untouched — these tests
deliberately target the Phase 4.1 fields (cache_read /
cache_write / reasoning / finish_reason / tool_names) that the
older tests didn't pin.

* fix(gate): forward real model + tools to /gate pre-flight (T4)

Pre-0.7.7 every SDK /gate call for any workflow with a budget was

hard-blocked because the runtime hard-coded the literal string

"budget-precheck" as the model. The backend's PolicyEvaluationGraph

treated any synthetic cost_limit rule with score > 0.8 as Block,

so the pricing lookup never landed on a real model and the rule

fired with the wrong score.

This commit:

* Adds nullrun.set_call_context(model=..., tools=[...]) plus

  get_call_model / get_call_tools helpers (and the underlying

  _call_model_var / _call_tools_var contextvars in

  nullrun.context).

* Wires the call context into check_workflow_budget: the /gate

  payload now carries the real model name (or None when unset)

  and the user-supplied tool list. tools=[] vs missing-None are

  distinguished on the wire per gate/internal.rs::check_tool_block.

* Transport.check forwards the tools key when set (it was

  silently dropped pre-fix).

* tests/conftest.py reset_runtime clears the new contextvars so

  a test's set_call_context(...) doesn't leak into the next

  test's wire payload.

* New tests/test_gate_real_path.py pins down the regression:

  default request allows a clean workflow, real block still

  honored, no policy-N residue on the wire, set_call_context

  flows into the body, no-context means no tools key, and the

  helpers are reachable from nullrun.*.

Bumps version to 0.7.7. No breaking changes - new helpers

default to None / empty so existing call sites keep working.
maltsev-dev added a commit that referenced this pull request Jun 28, 2026
* release: 0.7.6 — FastAPI integration + user-facing message catalog

Additive patch on top of the 0.7.0 thin-client refactor. No
breaking changes.

Added
-----

* nullrun.integrations.fastapi — one-line FastAPI integration
  that turns every NullRunDecision / NullRunInfrastructureError
  thrown by @nullrun.protect endpoints into a clean JSON
  response with the right HTTP status code. No per-endpoint
  except blocks required.

  Response shape:
    {"error_code": "NR-B004",
     "user_message": "You've reached the usage limit...",
     "category": "decision"}

  HTTP status mapping:
    * NR-B004 (budget), NR-L001 (loop), NR-R001 (rate) -> 429
      with optional Retry-After
    * NR-T001 (tool blocked), NR-X001 (generic block) -> 403
    * NR-W003 (paused) -> 503 with Retry-After
    * NR-W002 (killed) -> 503; WorkflowKilledInterrupt is a
      BaseException subclass so Starlette's
      add_exception_handler refuses it — handled via ASGI
      middleware instead (hybrid pattern, documented in
      module docstring).
    * NullRunInfrastructureError subclasses -> 503 (our side,
      not user's).

* nullrun.messages — default user-facing message catalog.
  Every NR-* error code has an English default message owned
  by NULLRUN, not customer code. Customer Support Bots hitting
  a budget cap show the same wording across every NullRun-backed
  application.
    * format_user_message(exc) — render exception as user-facing
      string
    * set_user_message(code, text) — per-process override for
      branded variants
    * get_user_message(code) — raw lookup
    * reset_overrides() — clear all overrides (for tests)

Changed
-------

* Transport._send_batch canonical JSON serialization — route the
  /track/batch body through _signed_request_body for consistent
  compact-separator serialization. HMAC itself is unaffected,
  but consistent serialization removes a special-case from the
  wire-format contract tests.

* Transport._send_batch actions response handling — backend
  renamed BatchTrackResponse.actions_taken (debug names) ->
  BatchTrackResponse.actions (ActionTaken structs). Read both
  for forward-compat; per-element try/except so one malformed
  entry doesn't abort the whole loop.

* pyproject.toml metadata — long-form description with search
  keywords, Maintainer: populated via maintainers=[...],
  expanded classifiers (Linux / Windows / macOS, Python 3.13,
  CPython, Security / AI / WWW/HTTP topics), project URL
  expander.

Tests
-----

* tests/test_messages.py (new, 282 lines) — catalog
  completeness (every NR-* code has a default message),
  override / reset behavior, render path.
* tests/test_integrations_fastapi.py (new, 289 lines) — HTTP
  status mapping per error code, response shape, ASGI
  middleware path for WorkflowKilledInterrupt, hybrid
  composition.
* tests/test_decision_split.py (new, 199 lines) — pins the
  decision / infrastructure error split.
* Updates to tests/test_runtime.py, tests/test_extractors.py
  reflecting transport canonical-JSON + actions-renamed
  changes.

Release plumbing
----------------

* pyproject.toml: version bumped 0.7.0 -> 0.7.6
* src/nullrun/__version__.py: __version__ = "0.7.6"
* CHANGELOG.md: full 0.7.6 entry covering additions,
  transport changes, metadata improvements

Tests pass locally (per session log) — pytest on Windows /
Python 3.14.2 is green.

* ci: fix PR #35 — fastapi dep + Transport._send_batch typo + coverage padding

PR #35 (release/0.7.6) failed all four CI jobs (test 3.10/3.11/3.12,
coverage, codecov/patch) on the same root cause + one latent bug
masked by it. This commit lands the fixes plus the last-mile tests
that bring coverage above the 82% threshold.

CI failure root
---------------

* tests/test_integrations_fastapi.py does from fastapi import ...
  at module top-level. CI installs only pip install -e '.[dev]',
  and fastapi was declared as an *optional* [fastapi] extra,
  NOT in [dev]. Pytest collection aborted with
  ModuleNotFoundError: No module named 'fastapi' → all 4 jobs red.
* Fix: add fastapi>=0.100,<1.0 to [dev]. Same precedent as
  langchain-core (already in [dev] for the same import-time
  contract: nullrun.instrumentation.langgraph is eager-imported
  from nullrun.decorators at collection time, so the test extras
  must cover the import chain).

Latent bug surfaced by the first fix
------------------------------------

The same PR refactored Transport._send_batch_with_retry_info to
route the /track/batch body through _signed_request_body for
canonical-JSON serialization (matching /gate and /execute). The two
sibling call sites use the module-level helper _signed_request_body
(no self.); this one used self._signed_request_body by typo.
Result: AttributeError on every batch flush, breaking 15 existing
tests across test_transport.py / test_track_batch_retry.py /
test_integration_contract.py / test_signal_safety.py. As long as
the fastapi collection error aborted pytest, this was hidden. Fixed
to _signed_request_body(...) with a docstring noting why it is
module-level and what the bug looked like.

Coverage padding (codecov/patch was failing on this too)
--------------------------------------------------------

Total coverage on the failing CI run was 81.98% — 0.02pp under the
fail-under=82 gate. After the two fixes above it would have
recovered to ~82.0% on the dot, so I added minimal tests for the
cheapest-to-cover gaps:

* tests/test_breaker_main.py (new) — covers the 5 statements in
  nullrun.breaker.__main__.main() (0% → 100%). The module
  exists so python -m nullrun.breaker exits cleanly instead of
  failing with No module named nullrun.breaker.__main__; the
  previous fix-mechanism was return 0 after a print, but no
  test was exercising it.
* tests/test_status.py — extends TestSummary with seven
  scenarios covering each conditional branch of NullRunStatus.summary()
  (organization_id, workflow_id, workflow_state != Normal,
  backend_reachable=False, ws_connected=False, recent_errors).
  status.py jumps 84.52% → 98.81%.
* tests/test_integrations_fastapi.py — four tests on
  _build_headers covering non-numeric, zero, negative, and
  resume_after (the WorkflowPausedException code path).
  integrations/fastapi.py jumps 90.22% → 94.57%.

After all three: TOTAL 81.98% → 82.46%, comfortably above the gate.

Verification
------------

* Local pytest: 997 passed, 13 skipped, 0 failed
  (Windows / Python 3.14.2, 8m47s — same env the original commit
  was validated in).
* python -m coverage report — 82.46%, no fail-under complaint.

* test: cover Phase 4.1 instrumentation — finish_reason + cache/reasoning/tools

Patch coverage on PR #35 was 62.38% against a 65% threshold (codecov
target 70% / threshold 5pp). The two biggest delta-holders against
master were auto.py (+286) and langgraph.py (+221), both dominated
by Phase 4.1 additions:

  * auto._normalize_finish_reason + _FINISH_REASON_MAP
  * auto._openai_extractor  second-tier fields (cache_read_tokens,
    cache_write_tokens, reasoning_tokens, finish_reason, tool_names)
  * auto._anthropic_extractor cache_read / cache_write
  * langgraph._safe_get_gen_message
  * langgraph._get_finish_reason (5-source fallback chain)
  * langgraph.extract_usage_from_response second-tier fields

These are pure / near-pure functions with no network or vendor SDK
calls. Coverage padding is cheap — pin the canonical wire shapes
once and the backend ingest contract gets a free live spec.

Local numbers:
  * auto.py        63.44% -> 64.01%   (file-level, +57 statements)
  * langgraph.py   78.50% -> 86.01%   (file-level, +32 statements)
  * TOTAL          82.46% -> 83.13%   (already above 82% gate)

41 tests, all green. Existing test_extractors.py and
test_langgraph_callback.py left untouched — these tests
deliberately target the Phase 4.1 fields (cache_read /
cache_write / reasoning / finish_reason / tool_names) that the
older tests didn't pin.

* fix(gate): forward real model + tools to /gate pre-flight (T4)

Pre-0.7.7 every SDK /gate call for any workflow with a budget was

hard-blocked because the runtime hard-coded the literal string

"budget-precheck" as the model. The backend's PolicyEvaluationGraph

treated any synthetic cost_limit rule with score > 0.8 as Block,

so the pricing lookup never landed on a real model and the rule

fired with the wrong score.

This commit:

* Adds nullrun.set_call_context(model=..., tools=[...]) plus

  get_call_model / get_call_tools helpers (and the underlying

  _call_model_var / _call_tools_var contextvars in

  nullrun.context).

* Wires the call context into check_workflow_budget: the /gate

  payload now carries the real model name (or None when unset)

  and the user-supplied tool list. tools=[] vs missing-None are

  distinguished on the wire per gate/internal.rs::check_tool_block.

* Transport.check forwards the tools key when set (it was

  silently dropped pre-fix).

* tests/conftest.py reset_runtime clears the new contextvars so

  a test's set_call_context(...) doesn't leak into the next

  test's wire payload.

* New tests/test_gate_real_path.py pins down the regression:

  default request allows a clean workflow, real block still

  honored, no policy-N residue on the wire, set_call_context

  flows into the body, no-context means no tools key, and the

  helpers are reachable from nullrun.*.

Bumps version to 0.7.7. No breaking changes - new helpers

default to None / empty so existing call sites keep working.

* release: 0.7.8 — fail-loud on deprecated surface

Two silent fail-OPEN footguns are converted to explicit
DeprecationWarning / RuntimeError so misconfigurations show up at
SDK init instead of being diagnosed from a missing proto trace.

Deprecated:

* NullRunRuntime.start_recording() and .stop_recording() now emit
  DeprecationWarning. They have been silent no-op stubs since
  Sprint 2.1 (0.4.0) — decision history is now on the backend
  dashboard at /control-center/decision-history. Both methods
  will be removed in 0.9.0.

* NULLRUN_USE_GRPC=1 now raises RuntimeError at SDK init instead
  of silently falling back to HTTP with an info log. gRPC is on
  the roadmap but not implemented; unset the env var to use HTTP.

Hardening (init path):

* Transport._post_auth_with_retry (new) — retry transient 503 / 504
  + network blips during /api/v1/auth/verify. Backend emits 503
  + Retry-After: 5 on transient DB errors (handlers.rs:11346-51).
  Pre-fix the first 503 surfaced as NR-A001 to the user as if the
  API key were bad. Three attempts, exponential backoff
  (0.5s → 1s → 2s), honors Retry-After when present. Auth-key
  failures (401) are NOT retried — a wrong key on attempt 1 is a
  wrong key on attempt 3.

Transport refactor:

* Transport._add_hmac_headers (new) — pulls the HMAC header
  construction out of _signed_request_body so /track/batch,
  /gate, /check, /execute all share one source of truth for
  Content-Type / X-Signature / X-Signature-Timestamp / X-API-Key
  / Authorization headers. HMAC formula unchanged.

* generate_hmac_signature + verify_hmac_signature accept str | bytes
  for body. Legacy str callers (and the FastAPI integration) keep
  working without an explicit .encode().

* actions_taken → actions on /track/batch response. Backend renamed
  BatchTrackResponse.actions_taken (debug names) → actions
  (ActionTaken structs with human-readable strings moved to
  messages). Read both keys for forward-compat.

Test updates:

* tests/test_framework_patches — alignment with retry + actions
  rename.
* tests/test_high_reliability_fixes — re-pinned for _post_auth_with_retry.
* tests/test_hmac_signing — expanded for str/bytes body + new
  _add_hmac_headers helper.
* tests/test_integration_contract — backend actions rename covered.
* tests/test_transport — retry semantics.

Bumps version to 0.7.8. No breaking changes for callers who don't
touch start_recording / stop_recording / NULLRUN_USE_GRPC.

* test(grpc): align test_grpc_removed with 0.7.8 NULLRUN_USE_GRPC contract

The 0.7.8 commit changed NULLRUN_USE_GRPC=1 from silent no-op +
INFO log to an explicit RuntimeError, but the regression test
in tests/test_grpc_removed.py still pinned the old behavior
(``test_nullrun_use_grpc_does_not_crash_init`` asserting
make_runtime() succeeded and an INFO line was logged).

CI on PR #38 failed on this test:

  FAILED tests/test_grpc_removed.py::TestGrpcRemoved
    ::test_nullrun_use_grpc_does_not_crash_init
  E   RuntimeError: NULLRUN_USE_GRPC is set but the gRPC
      transport is not yet implemented. ...

This commit updates the test to pin the new 0.7.8 contract:
the env var must raise RuntimeError, and the error message
must name the offending variable + point at the docs page.

The test is renamed from
``test_nullrun_use_grpc_does_not_crash_init`` to
``test_nullrun_use_grpc_raises_runtime_error`` so the test
name itself documents the new contract.

The module docstring (point 2 in the contract list) is
updated to say "raises RuntimeError" instead of "does NOT
crash init — it logs an INFO line and silently falls back
to HTTP". The 0.3.1 -> 0.7.8 evolution is documented in the
test docstring as a contract-evolution footnote for future
maintainers.

Imports: removed unused `import logging` and `caplog`
parameter (no longer asserting on log records); added
`import pytest` for `pytest.raises`.

No production-code change. No version bump. The fix is
self-contained to tests/test_grpc_removed.py.

* style(runtime): sort stdlib imports (ruff I001)

The 0.7.8 commit (fail-loud on deprecated surface) added
``import warnings`` mid-block in src/nullrun/runtime.py:34,
breaking alphabetical order:

    asyncio
    logging
    os
    warnings       <-- out of order
    threading
    time
    uuid

Ruff on PR #38 CI (Run ruff check src/) flagged it as I001.

Reorder to alphabetical:

    asyncio
    logging
    os
    threading
    time
    uuid
    warnings

Verified:
  * ruff check src/ -> All checks passed!
  * pytest tests/test_grpc_removed.py tests/test_runtime_branches.py
    -> 47 passed

No behavior change, no production logic touched. Pure lint fix.
maltsev-dev added a commit that referenced this pull request Jun 28, 2026
* release: 0.7.6 — FastAPI integration + user-facing message catalog

Additive patch on top of the 0.7.0 thin-client refactor. No
breaking changes.

Added
-----

* nullrun.integrations.fastapi — one-line FastAPI integration
  that turns every NullRunDecision / NullRunInfrastructureError
  thrown by @nullrun.protect endpoints into a clean JSON
  response with the right HTTP status code. No per-endpoint
  except blocks required.

  Response shape:
    {"error_code": "NR-B004",
     "user_message": "You've reached the usage limit...",
     "category": "decision"}

  HTTP status mapping:
    * NR-B004 (budget), NR-L001 (loop), NR-R001 (rate) -> 429
      with optional Retry-After
    * NR-T001 (tool blocked), NR-X001 (generic block) -> 403
    * NR-W003 (paused) -> 503 with Retry-After
    * NR-W002 (killed) -> 503; WorkflowKilledInterrupt is a
      BaseException subclass so Starlette's
      add_exception_handler refuses it — handled via ASGI
      middleware instead (hybrid pattern, documented in
      module docstring).
    * NullRunInfrastructureError subclasses -> 503 (our side,
      not user's).

* nullrun.messages — default user-facing message catalog.
  Every NR-* error code has an English default message owned
  by NULLRUN, not customer code. Customer Support Bots hitting
  a budget cap show the same wording across every NullRun-backed
  application.
    * format_user_message(exc) — render exception as user-facing
      string
    * set_user_message(code, text) — per-process override for
      branded variants
    * get_user_message(code) — raw lookup
    * reset_overrides() — clear all overrides (for tests)

Changed
-------

* Transport._send_batch canonical JSON serialization — route the
  /track/batch body through _signed_request_body for consistent
  compact-separator serialization. HMAC itself is unaffected,
  but consistent serialization removes a special-case from the
  wire-format contract tests.

* Transport._send_batch actions response handling — backend
  renamed BatchTrackResponse.actions_taken (debug names) ->
  BatchTrackResponse.actions (ActionTaken structs). Read both
  for forward-compat; per-element try/except so one malformed
  entry doesn't abort the whole loop.

* pyproject.toml metadata — long-form description with search
  keywords, Maintainer: populated via maintainers=[...],
  expanded classifiers (Linux / Windows / macOS, Python 3.13,
  CPython, Security / AI / WWW/HTTP topics), project URL
  expander.

Tests
-----

* tests/test_messages.py (new, 282 lines) — catalog
  completeness (every NR-* code has a default message),
  override / reset behavior, render path.
* tests/test_integrations_fastapi.py (new, 289 lines) — HTTP
  status mapping per error code, response shape, ASGI
  middleware path for WorkflowKilledInterrupt, hybrid
  composition.
* tests/test_decision_split.py (new, 199 lines) — pins the
  decision / infrastructure error split.
* Updates to tests/test_runtime.py, tests/test_extractors.py
  reflecting transport canonical-JSON + actions-renamed
  changes.

Release plumbing
----------------

* pyproject.toml: version bumped 0.7.0 -> 0.7.6
* src/nullrun/__version__.py: __version__ = "0.7.6"
* CHANGELOG.md: full 0.7.6 entry covering additions,
  transport changes, metadata improvements

Tests pass locally (per session log) — pytest on Windows /
Python 3.14.2 is green.

* ci: fix PR #35 — fastapi dep + Transport._send_batch typo + coverage padding

PR #35 (release/0.7.6) failed all four CI jobs (test 3.10/3.11/3.12,
coverage, codecov/patch) on the same root cause + one latent bug
masked by it. This commit lands the fixes plus the last-mile tests
that bring coverage above the 82% threshold.

CI failure root
---------------

* tests/test_integrations_fastapi.py does from fastapi import ...
  at module top-level. CI installs only pip install -e '.[dev]',
  and fastapi was declared as an *optional* [fastapi] extra,
  NOT in [dev]. Pytest collection aborted with
  ModuleNotFoundError: No module named 'fastapi' → all 4 jobs red.
* Fix: add fastapi>=0.100,<1.0 to [dev]. Same precedent as
  langchain-core (already in [dev] for the same import-time
  contract: nullrun.instrumentation.langgraph is eager-imported
  from nullrun.decorators at collection time, so the test extras
  must cover the import chain).

Latent bug surfaced by the first fix
------------------------------------

The same PR refactored Transport._send_batch_with_retry_info to
route the /track/batch body through _signed_request_body for
canonical-JSON serialization (matching /gate and /execute). The two
sibling call sites use the module-level helper _signed_request_body
(no self.); this one used self._signed_request_body by typo.
Result: AttributeError on every batch flush, breaking 15 existing
tests across test_transport.py / test_track_batch_retry.py /
test_integration_contract.py / test_signal_safety.py. As long as
the fastapi collection error aborted pytest, this was hidden. Fixed
to _signed_request_body(...) with a docstring noting why it is
module-level and what the bug looked like.

Coverage padding (codecov/patch was failing on this too)
--------------------------------------------------------

Total coverage on the failing CI run was 81.98% — 0.02pp under the
fail-under=82 gate. After the two fixes above it would have
recovered to ~82.0% on the dot, so I added minimal tests for the
cheapest-to-cover gaps:

* tests/test_breaker_main.py (new) — covers the 5 statements in
  nullrun.breaker.__main__.main() (0% → 100%). The module
  exists so python -m nullrun.breaker exits cleanly instead of
  failing with No module named nullrun.breaker.__main__; the
  previous fix-mechanism was return 0 after a print, but no
  test was exercising it.
* tests/test_status.py — extends TestSummary with seven
  scenarios covering each conditional branch of NullRunStatus.summary()
  (organization_id, workflow_id, workflow_state != Normal,
  backend_reachable=False, ws_connected=False, recent_errors).
  status.py jumps 84.52% → 98.81%.
* tests/test_integrations_fastapi.py — four tests on
  _build_headers covering non-numeric, zero, negative, and
  resume_after (the WorkflowPausedException code path).
  integrations/fastapi.py jumps 90.22% → 94.57%.

After all three: TOTAL 81.98% → 82.46%, comfortably above the gate.

Verification
------------

* Local pytest: 997 passed, 13 skipped, 0 failed
  (Windows / Python 3.14.2, 8m47s — same env the original commit
  was validated in).
* python -m coverage report — 82.46%, no fail-under complaint.

* test: cover Phase 4.1 instrumentation — finish_reason + cache/reasoning/tools

Patch coverage on PR #35 was 62.38% against a 65% threshold (codecov
target 70% / threshold 5pp). The two biggest delta-holders against
master were auto.py (+286) and langgraph.py (+221), both dominated
by Phase 4.1 additions:

  * auto._normalize_finish_reason + _FINISH_REASON_MAP
  * auto._openai_extractor  second-tier fields (cache_read_tokens,
    cache_write_tokens, reasoning_tokens, finish_reason, tool_names)
  * auto._anthropic_extractor cache_read / cache_write
  * langgraph._safe_get_gen_message
  * langgraph._get_finish_reason (5-source fallback chain)
  * langgraph.extract_usage_from_response second-tier fields

These are pure / near-pure functions with no network or vendor SDK
calls. Coverage padding is cheap — pin the canonical wire shapes
once and the backend ingest contract gets a free live spec.

Local numbers:
  * auto.py        63.44% -> 64.01%   (file-level, +57 statements)
  * langgraph.py   78.50% -> 86.01%   (file-level, +32 statements)
  * TOTAL          82.46% -> 83.13%   (already above 82% gate)

41 tests, all green. Existing test_extractors.py and
test_langgraph_callback.py left untouched — these tests
deliberately target the Phase 4.1 fields (cache_read /
cache_write / reasoning / finish_reason / tool_names) that the
older tests didn't pin.

* fix(gate): forward real model + tools to /gate pre-flight (T4)

Pre-0.7.7 every SDK /gate call for any workflow with a budget was

hard-blocked because the runtime hard-coded the literal string

"budget-precheck" as the model. The backend's PolicyEvaluationGraph

treated any synthetic cost_limit rule with score > 0.8 as Block,

so the pricing lookup never landed on a real model and the rule

fired with the wrong score.

This commit:

* Adds nullrun.set_call_context(model=..., tools=[...]) plus

  get_call_model / get_call_tools helpers (and the underlying

  _call_model_var / _call_tools_var contextvars in

  nullrun.context).

* Wires the call context into check_workflow_budget: the /gate

  payload now carries the real model name (or None when unset)

  and the user-supplied tool list. tools=[] vs missing-None are

  distinguished on the wire per gate/internal.rs::check_tool_block.

* Transport.check forwards the tools key when set (it was

  silently dropped pre-fix).

* tests/conftest.py reset_runtime clears the new contextvars so

  a test's set_call_context(...) doesn't leak into the next

  test's wire payload.

* New tests/test_gate_real_path.py pins down the regression:

  default request allows a clean workflow, real block still

  honored, no policy-N residue on the wire, set_call_context

  flows into the body, no-context means no tools key, and the

  helpers are reachable from nullrun.*.

Bumps version to 0.7.7. No breaking changes - new helpers

default to None / empty so existing call sites keep working.

* release: 0.7.8 — fail-loud on deprecated surface

Two silent fail-OPEN footguns are converted to explicit
DeprecationWarning / RuntimeError so misconfigurations show up at
SDK init instead of being diagnosed from a missing proto trace.

Deprecated:

* NullRunRuntime.start_recording() and .stop_recording() now emit
  DeprecationWarning. They have been silent no-op stubs since
  Sprint 2.1 (0.4.0) — decision history is now on the backend
  dashboard at /control-center/decision-history. Both methods
  will be removed in 0.9.0.

* NULLRUN_USE_GRPC=1 now raises RuntimeError at SDK init instead
  of silently falling back to HTTP with an info log. gRPC is on
  the roadmap but not implemented; unset the env var to use HTTP.

Hardening (init path):

* Transport._post_auth_with_retry (new) — retry transient 503 / 504
  + network blips during /api/v1/auth/verify. Backend emits 503
  + Retry-After: 5 on transient DB errors (handlers.rs:11346-51).
  Pre-fix the first 503 surfaced as NR-A001 to the user as if the
  API key were bad. Three attempts, exponential backoff
  (0.5s → 1s → 2s), honors Retry-After when present. Auth-key
  failures (401) are NOT retried — a wrong key on attempt 1 is a
  wrong key on attempt 3.

Transport refactor:

* Transport._add_hmac_headers (new) — pulls the HMAC header
  construction out of _signed_request_body so /track/batch,
  /gate, /check, /execute all share one source of truth for
  Content-Type / X-Signature / X-Signature-Timestamp / X-API-Key
  / Authorization headers. HMAC formula unchanged.

* generate_hmac_signature + verify_hmac_signature accept str | bytes
  for body. Legacy str callers (and the FastAPI integration) keep
  working without an explicit .encode().

* actions_taken → actions on /track/batch response. Backend renamed
  BatchTrackResponse.actions_taken (debug names) → actions
  (ActionTaken structs with human-readable strings moved to
  messages). Read both keys for forward-compat.

Test updates:

* tests/test_framework_patches — alignment with retry + actions
  rename.
* tests/test_high_reliability_fixes — re-pinned for _post_auth_with_retry.
* tests/test_hmac_signing — expanded for str/bytes body + new
  _add_hmac_headers helper.
* tests/test_integration_contract — backend actions rename covered.
* tests/test_transport — retry semantics.

Bumps version to 0.7.8. No breaking changes for callers who don't
touch start_recording / stop_recording / NULLRUN_USE_GRPC.

* test(grpc): align test_grpc_removed with 0.7.8 NULLRUN_USE_GRPC contract

The 0.7.8 commit changed NULLRUN_USE_GRPC=1 from silent no-op +
INFO log to an explicit RuntimeError, but the regression test
in tests/test_grpc_removed.py still pinned the old behavior
(``test_nullrun_use_grpc_does_not_crash_init`` asserting
make_runtime() succeeded and an INFO line was logged).

CI on PR #38 failed on this test:

  FAILED tests/test_grpc_removed.py::TestGrpcRemoved
    ::test_nullrun_use_grpc_does_not_crash_init
  E   RuntimeError: NULLRUN_USE_GRPC is set but the gRPC
      transport is not yet implemented. ...

This commit updates the test to pin the new 0.7.8 contract:
the env var must raise RuntimeError, and the error message
must name the offending variable + point at the docs page.

The test is renamed from
``test_nullrun_use_grpc_does_not_crash_init`` to
``test_nullrun_use_grpc_raises_runtime_error`` so the test
name itself documents the new contract.

The module docstring (point 2 in the contract list) is
updated to say "raises RuntimeError" instead of "does NOT
crash init — it logs an INFO line and silently falls back
to HTTP". The 0.3.1 -> 0.7.8 evolution is documented in the
test docstring as a contract-evolution footnote for future
maintainers.

Imports: removed unused `import logging` and `caplog`
parameter (no longer asserting on log records); added
`import pytest` for `pytest.raises`.

No production-code change. No version bump. The fix is
self-contained to tests/test_grpc_removed.py.

* style(runtime): sort stdlib imports (ruff I001)

The 0.7.8 commit (fail-loud on deprecated surface) added
``import warnings`` mid-block in src/nullrun/runtime.py:34,
breaking alphabetical order:

    asyncio
    logging
    os
    warnings       <-- out of order
    threading
    time
    uuid

Ruff on PR #38 CI (Run ruff check src/) flagged it as I001.

Reorder to alphabetical:

    asyncio
    logging
    os
    threading
    time
    uuid
    warnings

Verified:
  * ruff check src/ -> All checks passed!
  * pytest tests/test_grpc_removed.py tests/test_runtime_branches.py
    -> 47 passed

No behavior change, no production logic touched. Pure lint fix.

* release: 0.8.0 — SDK wire-format audit (model/provider extraction)

Closes a class of silent-fail-OPEN path that was sending
model=None or model="unknown" on /track for many LLM-vendor
paths. Every such event cost the backend a model_pricing
lookup that returned no row, fell through to DEFAULT_RATE
(~$30/M), and emitted a fallback warning the operator
couldn't reproduce because the offending observation was
buried in another package's telemetry.

No public-API break. No behavior change for callers whose
instrumentation already populates model correctly. Pure
wire-payload hygiene.

runtime.py — track():

* Strips None values from the wire payload (pre-0.8.0
  forwarded every key except _WIRE_STRIP_FIELDS, including
  keys whose value was None). Putting {"model": null} on
  the wire triggered backend unwrap_or("default") and a
  fallback warning. Dropping None keeps the diagnostic
  signal loud (the new WARN below fires on missing-key,
  which is what we want operators to see) instead of
  silent (the JSON-null case).

* Adds logger.warning("track(): llm_call event missing
  'model' field — backend will fall back to DEFAULT_RATE.
  event=...") — the single signal an operator needs to
  reproduce "which observation produced an llm_call
  without model set". Activated only for llm_call; other
  event types are silent.

instrumentation/langgraph.py — NullRunCallback.on_llm_end:

* New _extract_model_from_response + _extract_provider_from_response
  helpers (mirror _get_finish_reason's best-effort
  pattern). Fallback chain: invocation_params → response
  metadata → AIMessage response_metadata → llm_output →
  direct attribute. "unknown" is now a true last resort,
  not the common case.

instrumentation/llama_index.py:

* extract_from_event fallback chain: event.response.model
  → event.response.raw.model → usage['model']. Mock
  providers and adapter-style ChatResponse now ship a
  real model id.

instrumentation/autogen.py:

* on_messages fallback chain: self.model → result.model.
  OpenAI's response carries the actual model id (may
  differ from request if the server resolved an alias).

instrumentation/auto.py — _emit_from_span (openai-agents):

* span model fallback chain: span['model'] →
  usage['model'] → span['response_metadata']['model_name'].
  Some custom tracer configs leave span['model'] empty;
  the other two sources usually have it.

  Sets model on the event only when we have a real value
  (empty/None is dropped — relies on the new None-strip
  in track() to keep the operator warning loud).

Bumps version to 0.8.0. No breaking changes for callers
who don't touch the wire path directly.
maltsev-dev added a commit that referenced this pull request Jun 29, 2026
…40)

* release: 0.7.6 — FastAPI integration + user-facing message catalog

Additive patch on top of the 0.7.0 thin-client refactor. No
breaking changes.

Added
-----

* nullrun.integrations.fastapi — one-line FastAPI integration
  that turns every NullRunDecision / NullRunInfrastructureError
  thrown by @nullrun.protect endpoints into a clean JSON
  response with the right HTTP status code. No per-endpoint
  except blocks required.

  Response shape:
    {"error_code": "NR-B004",
     "user_message": "You've reached the usage limit...",
     "category": "decision"}

  HTTP status mapping:
    * NR-B004 (budget), NR-L001 (loop), NR-R001 (rate) -> 429
      with optional Retry-After
    * NR-T001 (tool blocked), NR-X001 (generic block) -> 403
    * NR-W003 (paused) -> 503 with Retry-After
    * NR-W002 (killed) -> 503; WorkflowKilledInterrupt is a
      BaseException subclass so Starlette's
      add_exception_handler refuses it — handled via ASGI
      middleware instead (hybrid pattern, documented in
      module docstring).
    * NullRunInfrastructureError subclasses -> 503 (our side,
      not user's).

* nullrun.messages — default user-facing message catalog.
  Every NR-* error code has an English default message owned
  by NULLRUN, not customer code. Customer Support Bots hitting
  a budget cap show the same wording across every NullRun-backed
  application.
    * format_user_message(exc) — render exception as user-facing
      string
    * set_user_message(code, text) — per-process override for
      branded variants
    * get_user_message(code) — raw lookup
    * reset_overrides() — clear all overrides (for tests)

Changed
-------

* Transport._send_batch canonical JSON serialization — route the
  /track/batch body through _signed_request_body for consistent
  compact-separator serialization. HMAC itself is unaffected,
  but consistent serialization removes a special-case from the
  wire-format contract tests.

* Transport._send_batch actions response handling — backend
  renamed BatchTrackResponse.actions_taken (debug names) ->
  BatchTrackResponse.actions (ActionTaken structs). Read both
  for forward-compat; per-element try/except so one malformed
  entry doesn't abort the whole loop.

* pyproject.toml metadata — long-form description with search
  keywords, Maintainer: populated via maintainers=[...],
  expanded classifiers (Linux / Windows / macOS, Python 3.13,
  CPython, Security / AI / WWW/HTTP topics), project URL
  expander.

Tests
-----

* tests/test_messages.py (new, 282 lines) — catalog
  completeness (every NR-* code has a default message),
  override / reset behavior, render path.
* tests/test_integrations_fastapi.py (new, 289 lines) — HTTP
  status mapping per error code, response shape, ASGI
  middleware path for WorkflowKilledInterrupt, hybrid
  composition.
* tests/test_decision_split.py (new, 199 lines) — pins the
  decision / infrastructure error split.
* Updates to tests/test_runtime.py, tests/test_extractors.py
  reflecting transport canonical-JSON + actions-renamed
  changes.

Release plumbing
----------------

* pyproject.toml: version bumped 0.7.0 -> 0.7.6
* src/nullrun/__version__.py: __version__ = "0.7.6"
* CHANGELOG.md: full 0.7.6 entry covering additions,
  transport changes, metadata improvements

Tests pass locally (per session log) — pytest on Windows /
Python 3.14.2 is green.

* ci: fix PR #35 — fastapi dep + Transport._send_batch typo + coverage padding

PR #35 (release/0.7.6) failed all four CI jobs (test 3.10/3.11/3.12,
coverage, codecov/patch) on the same root cause + one latent bug
masked by it. This commit lands the fixes plus the last-mile tests
that bring coverage above the 82% threshold.

CI failure root
---------------

* tests/test_integrations_fastapi.py does from fastapi import ...
  at module top-level. CI installs only pip install -e '.[dev]',
  and fastapi was declared as an *optional* [fastapi] extra,
  NOT in [dev]. Pytest collection aborted with
  ModuleNotFoundError: No module named 'fastapi' → all 4 jobs red.
* Fix: add fastapi>=0.100,<1.0 to [dev]. Same precedent as
  langchain-core (already in [dev] for the same import-time
  contract: nullrun.instrumentation.langgraph is eager-imported
  from nullrun.decorators at collection time, so the test extras
  must cover the import chain).

Latent bug surfaced by the first fix
------------------------------------

The same PR refactored Transport._send_batch_with_retry_info to
route the /track/batch body through _signed_request_body for
canonical-JSON serialization (matching /gate and /execute). The two
sibling call sites use the module-level helper _signed_request_body
(no self.); this one used self._signed_request_body by typo.
Result: AttributeError on every batch flush, breaking 15 existing
tests across test_transport.py / test_track_batch_retry.py /
test_integration_contract.py / test_signal_safety.py. As long as
the fastapi collection error aborted pytest, this was hidden. Fixed
to _signed_request_body(...) with a docstring noting why it is
module-level and what the bug looked like.

Coverage padding (codecov/patch was failing on this too)
--------------------------------------------------------

Total coverage on the failing CI run was 81.98% — 0.02pp under the
fail-under=82 gate. After the two fixes above it would have
recovered to ~82.0% on the dot, so I added minimal tests for the
cheapest-to-cover gaps:

* tests/test_breaker_main.py (new) — covers the 5 statements in
  nullrun.breaker.__main__.main() (0% → 100%). The module
  exists so python -m nullrun.breaker exits cleanly instead of
  failing with No module named nullrun.breaker.__main__; the
  previous fix-mechanism was return 0 after a print, but no
  test was exercising it.
* tests/test_status.py — extends TestSummary with seven
  scenarios covering each conditional branch of NullRunStatus.summary()
  (organization_id, workflow_id, workflow_state != Normal,
  backend_reachable=False, ws_connected=False, recent_errors).
  status.py jumps 84.52% → 98.81%.
* tests/test_integrations_fastapi.py — four tests on
  _build_headers covering non-numeric, zero, negative, and
  resume_after (the WorkflowPausedException code path).
  integrations/fastapi.py jumps 90.22% → 94.57%.

After all three: TOTAL 81.98% → 82.46%, comfortably above the gate.

Verification
------------

* Local pytest: 997 passed, 13 skipped, 0 failed
  (Windows / Python 3.14.2, 8m47s — same env the original commit
  was validated in).
* python -m coverage report — 82.46%, no fail-under complaint.

* test: cover Phase 4.1 instrumentation — finish_reason + cache/reasoning/tools

Patch coverage on PR #35 was 62.38% against a 65% threshold (codecov
target 70% / threshold 5pp). The two biggest delta-holders against
master were auto.py (+286) and langgraph.py (+221), both dominated
by Phase 4.1 additions:

  * auto._normalize_finish_reason + _FINISH_REASON_MAP
  * auto._openai_extractor  second-tier fields (cache_read_tokens,
    cache_write_tokens, reasoning_tokens, finish_reason, tool_names)
  * auto._anthropic_extractor cache_read / cache_write
  * langgraph._safe_get_gen_message
  * langgraph._get_finish_reason (5-source fallback chain)
  * langgraph.extract_usage_from_response second-tier fields

These are pure / near-pure functions with no network or vendor SDK
calls. Coverage padding is cheap — pin the canonical wire shapes
once and the backend ingest contract gets a free live spec.

Local numbers:
  * auto.py        63.44% -> 64.01%   (file-level, +57 statements)
  * langgraph.py   78.50% -> 86.01%   (file-level, +32 statements)
  * TOTAL          82.46% -> 83.13%   (already above 82% gate)

41 tests, all green. Existing test_extractors.py and
test_langgraph_callback.py left untouched — these tests
deliberately target the Phase 4.1 fields (cache_read /
cache_write / reasoning / finish_reason / tool_names) that the
older tests didn't pin.

* fix(gate): forward real model + tools to /gate pre-flight (T4)

Pre-0.7.7 every SDK /gate call for any workflow with a budget was

hard-blocked because the runtime hard-coded the literal string

"budget-precheck" as the model. The backend's PolicyEvaluationGraph

treated any synthetic cost_limit rule with score > 0.8 as Block,

so the pricing lookup never landed on a real model and the rule

fired with the wrong score.

This commit:

* Adds nullrun.set_call_context(model=..., tools=[...]) plus

  get_call_model / get_call_tools helpers (and the underlying

  _call_model_var / _call_tools_var contextvars in

  nullrun.context).

* Wires the call context into check_workflow_budget: the /gate

  payload now carries the real model name (or None when unset)

  and the user-supplied tool list. tools=[] vs missing-None are

  distinguished on the wire per gate/internal.rs::check_tool_block.

* Transport.check forwards the tools key when set (it was

  silently dropped pre-fix).

* tests/conftest.py reset_runtime clears the new contextvars so

  a test's set_call_context(...) doesn't leak into the next

  test's wire payload.

* New tests/test_gate_real_path.py pins down the regression:

  default request allows a clean workflow, real block still

  honored, no policy-N residue on the wire, set_call_context

  flows into the body, no-context means no tools key, and the

  helpers are reachable from nullrun.*.

Bumps version to 0.7.7. No breaking changes - new helpers

default to None / empty so existing call sites keep working.

* release: 0.7.8 — fail-loud on deprecated surface

Two silent fail-OPEN footguns are converted to explicit
DeprecationWarning / RuntimeError so misconfigurations show up at
SDK init instead of being diagnosed from a missing proto trace.

Deprecated:

* NullRunRuntime.start_recording() and .stop_recording() now emit
  DeprecationWarning. They have been silent no-op stubs since
  Sprint 2.1 (0.4.0) — decision history is now on the backend
  dashboard at /control-center/decision-history. Both methods
  will be removed in 0.9.0.

* NULLRUN_USE_GRPC=1 now raises RuntimeError at SDK init instead
  of silently falling back to HTTP with an info log. gRPC is on
  the roadmap but not implemented; unset the env var to use HTTP.

Hardening (init path):

* Transport._post_auth_with_retry (new) — retry transient 503 / 504
  + network blips during /api/v1/auth/verify. Backend emits 503
  + Retry-After: 5 on transient DB errors (handlers.rs:11346-51).
  Pre-fix the first 503 surfaced as NR-A001 to the user as if the
  API key were bad. Three attempts, exponential backoff
  (0.5s → 1s → 2s), honors Retry-After when present. Auth-key
  failures (401) are NOT retried — a wrong key on attempt 1 is a
  wrong key on attempt 3.

Transport refactor:

* Transport._add_hmac_headers (new) — pulls the HMAC header
  construction out of _signed_request_body so /track/batch,
  /gate, /check, /execute all share one source of truth for
  Content-Type / X-Signature / X-Signature-Timestamp / X-API-Key
  / Authorization headers. HMAC formula unchanged.

* generate_hmac_signature + verify_hmac_signature accept str | bytes
  for body. Legacy str callers (and the FastAPI integration) keep
  working without an explicit .encode().

* actions_taken → actions on /track/batch response. Backend renamed
  BatchTrackResponse.actions_taken (debug names) → actions
  (ActionTaken structs with human-readable strings moved to
  messages). Read both keys for forward-compat.

Test updates:

* tests/test_framework_patches — alignment with retry + actions
  rename.
* tests/test_high_reliability_fixes — re-pinned for _post_auth_with_retry.
* tests/test_hmac_signing — expanded for str/bytes body + new
  _add_hmac_headers helper.
* tests/test_integration_contract — backend actions rename covered.
* tests/test_transport — retry semantics.

Bumps version to 0.7.8. No breaking changes for callers who don't
touch start_recording / stop_recording / NULLRUN_USE_GRPC.

* test(grpc): align test_grpc_removed with 0.7.8 NULLRUN_USE_GRPC contract

The 0.7.8 commit changed NULLRUN_USE_GRPC=1 from silent no-op +
INFO log to an explicit RuntimeError, but the regression test
in tests/test_grpc_removed.py still pinned the old behavior
(``test_nullrun_use_grpc_does_not_crash_init`` asserting
make_runtime() succeeded and an INFO line was logged).

CI on PR #38 failed on this test:

  FAILED tests/test_grpc_removed.py::TestGrpcRemoved
    ::test_nullrun_use_grpc_does_not_crash_init
  E   RuntimeError: NULLRUN_USE_GRPC is set but the gRPC
      transport is not yet implemented. ...

This commit updates the test to pin the new 0.7.8 contract:
the env var must raise RuntimeError, and the error message
must name the offending variable + point at the docs page.

The test is renamed from
``test_nullrun_use_grpc_does_not_crash_init`` to
``test_nullrun_use_grpc_raises_runtime_error`` so the test
name itself documents the new contract.

The module docstring (point 2 in the contract list) is
updated to say "raises RuntimeError" instead of "does NOT
crash init — it logs an INFO line and silently falls back
to HTTP". The 0.3.1 -> 0.7.8 evolution is documented in the
test docstring as a contract-evolution footnote for future
maintainers.

Imports: removed unused `import logging` and `caplog`
parameter (no longer asserting on log records); added
`import pytest` for `pytest.raises`.

No production-code change. No version bump. The fix is
self-contained to tests/test_grpc_removed.py.

* style(runtime): sort stdlib imports (ruff I001)

The 0.7.8 commit (fail-loud on deprecated surface) added
``import warnings`` mid-block in src/nullrun/runtime.py:34,
breaking alphabetical order:

    asyncio
    logging
    os
    warnings       <-- out of order
    threading
    time
    uuid

Ruff on PR #38 CI (Run ruff check src/) flagged it as I001.

Reorder to alphabetical:

    asyncio
    logging
    os
    threading
    time
    uuid
    warnings

Verified:
  * ruff check src/ -> All checks passed!
  * pytest tests/test_grpc_removed.py tests/test_runtime_branches.py
    -> 47 passed

No behavior change, no production logic touched. Pure lint fix.

* release: 0.8.0 — SDK wire-format audit (model/provider extraction)

Closes a class of silent-fail-OPEN path that was sending
model=None or model="unknown" on /track for many LLM-vendor
paths. Every such event cost the backend a model_pricing
lookup that returned no row, fell through to DEFAULT_RATE
(~$30/M), and emitted a fallback warning the operator
couldn't reproduce because the offending observation was
buried in another package's telemetry.

No public-API break. No behavior change for callers whose
instrumentation already populates model correctly. Pure
wire-payload hygiene.

runtime.py — track():

* Strips None values from the wire payload (pre-0.8.0
  forwarded every key except _WIRE_STRIP_FIELDS, including
  keys whose value was None). Putting {"model": null} on
  the wire triggered backend unwrap_or("default") and a
  fallback warning. Dropping None keeps the diagnostic
  signal loud (the new WARN below fires on missing-key,
  which is what we want operators to see) instead of
  silent (the JSON-null case).

* Adds logger.warning("track(): llm_call event missing
  'model' field — backend will fall back to DEFAULT_RATE.
  event=...") — the single signal an operator needs to
  reproduce "which observation produced an llm_call
  without model set". Activated only for llm_call; other
  event types are silent.

instrumentation/langgraph.py — NullRunCallback.on_llm_end:

* New _extract_model_from_response + _extract_provider_from_response
  helpers (mirror _get_finish_reason's best-effort
  pattern). Fallback chain: invocation_params → response
  metadata → AIMessage response_metadata → llm_output →
  direct attribute. "unknown" is now a true last resort,
  not the common case.

instrumentation/llama_index.py:

* extract_from_event fallback chain: event.response.model
  → event.response.raw.model → usage['model']. Mock
  providers and adapter-style ChatResponse now ship a
  real model id.

instrumentation/autogen.py:

* on_messages fallback chain: self.model → result.model.
  OpenAI's response carries the actual model id (may
  differ from request if the server resolved an alias).

instrumentation/auto.py — _emit_from_span (openai-agents):

* span model fallback chain: span['model'] →
  usage['model'] → span['response_metadata']['model_name'].
  Some custom tracer configs leave span['model'] empty;
  the other two sources usually have it.

  Sets model on the event only when we have a real value
  (empty/None is dropped — relies on the new None-strip
  in track() to keep the operator warning loud).

Bumps version to 0.8.0. No breaking changes for callers
who don't touch the wire path directly.

* fix: 0.8.2 — coverage wire-shape (metadata nesting) + model fallback

Two coordinated fixes from the 0.8.0 wire-format audit:

1. Coverage counters under metadata
   - src/nullrun/runtime.py: track_coverage() emits seen/tracked/
     streaming_skipped dicts under event.metadata instead of the
     top level. SdkTrackRequest uses explicit fields with no
     #[serde(flatten)] catchall, so top-level keys were silently
     dropped by serde and the dashboard's last_coverage_pct was
     permanently null.
   - tests/test_coverage_report.py: pin the wire shape (regression
     test).

2. Model name extraction fallback (Issue 2)
   - src/nullrun/instrumentation/auto.py: when the response body
     extractor returns None for model (OpenAI Responses API,
     Anthropic streaming edge cases), fall back to the model
     string the user embedded in the request body via
     ChatOpenAI(model='gpt-4.1-mini'). Without this, every such
     call was zero-billed (backend unwrap_or('default') +
     DEFAULT_RATE ≈ $0/call).
   - tests/test_model_fallback.py: unit-test the helper.

3. Backend batch response schema contract tests
   - tests/test_batch_response_parsing.py: pin the post-2026-06-27
     BatchTrackResponse shape (actions: Vec<ActionTaken>,
     messages: Vec<String>) and document that the legacy
     actions_taken: Vec<String> field is intentionally dropped in
     0.8.0.
maltsev-dev added a commit that referenced this pull request Aug 7, 2026
Closes the P0/P1/P2/P3 issues from the security review (plan §10/§11.4).

Security / PCI-DSS / GDPR

- P0-1: Mask positional PII in `_enforce_sensitive_tool` by introspecting
  the wrapped function's signature and applying `SENSITIVE_ARG_KEYS` to
  positional params. Pre-fix, `charge("4111-…-1111", 50)` forwarded the
  PAN into `/execute` and the audit log.
- P0-6 / P3-3: `_safe_repr` now redacts BEFORE truncating. The pre-fix
  order truncated first, so `details={…}` past position 50 leaked
  verbatim. `_safe_repr` is now the single source of truth for the
  redact-then-truncate flow.

Cost-audit / reliability

- P0-3: Bounded chunked reads on the sync + async httpx transports
  (`MAX_RESPONSE_BYTES`, default 16 MiB, `NULLRUN_MAX_RESPONSE_BYTES`
  env override). Above the cap, tracking is skipped and
  `_coverage_streaming_skipped` is incremented. Replaces the
  `response.read()` / `await response.aread()` unbounded buffer that
  held entire LLM streaming bodies in memory.
- P0-4: `_do_flush_locked` re-queue on CB OPEN now drops the NEWEST
  non-critical events instead of the oldest. The oldest events
  (incident start, billing-period start) are exactly what a billing
  investigator needs; losing them silently broke monthly rollups.
  Control-plane events (`state_change`, `kill_received`,
  `policy_invalidated`, `key_rotated`) are preserved unconditionally
  so the dashboard KILL switch lands even under sustained backend
  outage.

Identity

- S-8 / P2-4: `agent()` now emits `str(uuid.uuid4())` (with dashes).
  Pre-fix the format was `f"agent-{uuid.uuid4().hex}"` — 32 hex chars,
  no dashes — and backend UUID-typed columns dropped these to NULL
  on insert. User-supplied names are still preserved verbatim.
- §7.2 #16: `workflow()` context manager now resets `span_id` (not
  only `workflow_id` / `trace_id`) so nested `with span()` blocks
  don't leave the inner span_id visible inside the workflow scope.

Resource leaks

- S-9: `_active_runs` on `NullRunCallback` is now an `OrderedDict`
  capped at 4096 with FIFO eviction. Pre-fix the dict grew
  unbounded when `on_chain_end` did not fire (some LangChain
  versions short-circuit the end hook on chain-body errors).
- S-10: WebSocket reconnect loop is now capped at 10 consecutive
  failures, then falls back to HTTP-poll. Pre-fix the loop ran
  forever when the backend was permanently down, leaking the
  WS thread.

Transport

- §7.2 #6: Separate `hmac_verify_expired_total` counter so SRE can
  distinguish clock-skew (NTP drift) from forged packets. Mirrored
  in both the HTTP and WebSocket verify paths.
- §7.2 #35: `CircuitBreaker.call` now dispatches the OPEN→HALF_OPEN
  jitter through `_maybe_apply_open_jitter_sync` /
  `_maybe_apply_open_jitter_async`. Pre-fix the jitter used
  `time.sleep` before dispatching to async, which blocked the
  caller's event loop on every transition.
- P2-1: `_coverage_seen` now bumps in the httpx path (sync + async).
  Pre-fix the counter was only bumped by the `requests` transport,
  so the dashboard's coverage view was empty for the dominant
  OpenAI / Anthropic / Gemini / Mistral / Cohere traffic.
- P2-3: `is_sensitive_tool` match is case-insensitive. Pre-fix
  `"stripe.charge"` did not match `"Stripe.Charge"`, bypassing the
  sensitive gate.

Concurrency

- §7.2 #39: New `_tools_lock` guards every mutation of
  `_strict_mode_tools` / `_sensitive_tools`. Same lock guards the
  coverage-counter bump+prune sequence (§7.2 #33) so two threads
  can't both observe the dict at length 4095 and both grow it to
  4097 before either prune lands.
- §7.2 #47: New `_langchain_lock` / `_langgraph_lock` guard the
  patch sequences end-to-end. Pre-fix two threads racing through
  `auto_instrument` could both pass the early `_x_patched` check
  and double-wrap `BaseCallbackManager` / `Pregel`.
- §7.2 #33: `_COVERAGE_CAP` (4096) bounds the per-host coverage
  dicts.

Webhook delivery

- P3-2: Exponential backoff (0.5s, 1s, 2s, 4s, 8s, 16s, 30s cap)
  replaces the previous linear schedule. Linear didn't back off
  fast enough under sustained outage — each KILL/PAUSE spawned
  its own delivery thread, producing 1000+ spinning threads
  hammering the dead endpoint.

WAL crash-recovery

- P1-5b: Atomic WAL writes (tmp + `fsync` + `os.replace`), 64 MiB
  rotation with `os.replace(wal, wal.1)`, replay drains both
  `wal.1` and `wal`. New `NULLRUN_WAL_PATH` / `NULLRUN_WAL_MAX_BYTES`
  env overrides for containers with `readOnlyRootFilesystem: true`.

Tests

8 new regression test files (57 tests total):
  test_agent_id_uuid.py, test_args_pii_masked.py,
  test_streaming_oom_cap.py, test_lru_active_runs.py,
  test_reconnect_cap.py, test_coverage_seen_httpx.py,
  test_webhook_backoff.py, test_redact.py

`test_buffer_invariants.py` extended with drop-newest +
critical-event preservation cases. `test_release_polish.py`
updated to pin the 5s cap on both the sync and async jitter
helpers (post §7.2 #35 split).

Full incident write-ups in CHANGELOG.md under the same P0/S/P tags.
maltsev-dev added a commit that referenced this pull request Aug 7, 2026
* fix: P0 security/stability hardening bundle

Closes the P0/P1/P2/P3 issues from the security review (plan §10/§11.4).

Security / PCI-DSS / GDPR

- P0-1: Mask positional PII in `_enforce_sensitive_tool` by introspecting
  the wrapped function's signature and applying `SENSITIVE_ARG_KEYS` to
  positional params. Pre-fix, `charge("4111-…-1111", 50)` forwarded the
  PAN into `/execute` and the audit log.
- P0-6 / P3-3: `_safe_repr` now redacts BEFORE truncating. The pre-fix
  order truncated first, so `details={…}` past position 50 leaked
  verbatim. `_safe_repr` is now the single source of truth for the
  redact-then-truncate flow.

Cost-audit / reliability

- P0-3: Bounded chunked reads on the sync + async httpx transports
  (`MAX_RESPONSE_BYTES`, default 16 MiB, `NULLRUN_MAX_RESPONSE_BYTES`
  env override). Above the cap, tracking is skipped and
  `_coverage_streaming_skipped` is incremented. Replaces the
  `response.read()` / `await response.aread()` unbounded buffer that
  held entire LLM streaming bodies in memory.
- P0-4: `_do_flush_locked` re-queue on CB OPEN now drops the NEWEST
  non-critical events instead of the oldest. The oldest events
  (incident start, billing-period start) are exactly what a billing
  investigator needs; losing them silently broke monthly rollups.
  Control-plane events (`state_change`, `kill_received`,
  `policy_invalidated`, `key_rotated`) are preserved unconditionally
  so the dashboard KILL switch lands even under sustained backend
  outage.

Identity

- S-8 / P2-4: `agent()` now emits `str(uuid.uuid4())` (with dashes).
  Pre-fix the format was `f"agent-{uuid.uuid4().hex}"` — 32 hex chars,
  no dashes — and backend UUID-typed columns dropped these to NULL
  on insert. User-supplied names are still preserved verbatim.
- §7.2 #16: `workflow()` context manager now resets `span_id` (not
  only `workflow_id` / `trace_id`) so nested `with span()` blocks
  don't leave the inner span_id visible inside the workflow scope.

Resource leaks

- S-9: `_active_runs` on `NullRunCallback` is now an `OrderedDict`
  capped at 4096 with FIFO eviction. Pre-fix the dict grew
  unbounded when `on_chain_end` did not fire (some LangChain
  versions short-circuit the end hook on chain-body errors).
- S-10: WebSocket reconnect loop is now capped at 10 consecutive
  failures, then falls back to HTTP-poll. Pre-fix the loop ran
  forever when the backend was permanently down, leaking the
  WS thread.

Transport

- §7.2 #6: Separate `hmac_verify_expired_total` counter so SRE can
  distinguish clock-skew (NTP drift) from forged packets. Mirrored
  in both the HTTP and WebSocket verify paths.
- §7.2 #35: `CircuitBreaker.call` now dispatches the OPEN→HALF_OPEN
  jitter through `_maybe_apply_open_jitter_sync` /
  `_maybe_apply_open_jitter_async`. Pre-fix the jitter used
  `time.sleep` before dispatching to async, which blocked the
  caller's event loop on every transition.
- P2-1: `_coverage_seen` now bumps in the httpx path (sync + async).
  Pre-fix the counter was only bumped by the `requests` transport,
  so the dashboard's coverage view was empty for the dominant
  OpenAI / Anthropic / Gemini / Mistral / Cohere traffic.
- P2-3: `is_sensitive_tool` match is case-insensitive. Pre-fix
  `"stripe.charge"` did not match `"Stripe.Charge"`, bypassing the
  sensitive gate.

Concurrency

- §7.2 #39: New `_tools_lock` guards every mutation of
  `_strict_mode_tools` / `_sensitive_tools`. Same lock guards the
  coverage-counter bump+prune sequence (§7.2 #33) so two threads
  can't both observe the dict at length 4095 and both grow it to
  4097 before either prune lands.
- §7.2 #47: New `_langchain_lock` / `_langgraph_lock` guard the
  patch sequences end-to-end. Pre-fix two threads racing through
  `auto_instrument` could both pass the early `_x_patched` check
  and double-wrap `BaseCallbackManager` / `Pregel`.
- §7.2 #33: `_COVERAGE_CAP` (4096) bounds the per-host coverage
  dicts.

Webhook delivery

- P3-2: Exponential backoff (0.5s, 1s, 2s, 4s, 8s, 16s, 30s cap)
  replaces the previous linear schedule. Linear didn't back off
  fast enough under sustained outage — each KILL/PAUSE spawned
  its own delivery thread, producing 1000+ spinning threads
  hammering the dead endpoint.

WAL crash-recovery

- P1-5b: Atomic WAL writes (tmp + `fsync` + `os.replace`), 64 MiB
  rotation with `os.replace(wal, wal.1)`, replay drains both
  `wal.1` and `wal`. New `NULLRUN_WAL_PATH` / `NULLRUN_WAL_MAX_BYTES`
  env overrides for containers with `readOnlyRootFilesystem: true`.

Tests

8 new regression test files (57 tests total):
  test_agent_id_uuid.py, test_args_pii_masked.py,
  test_streaming_oom_cap.py, test_lru_active_runs.py,
  test_reconnect_cap.py, test_coverage_seen_httpx.py,
  test_webhook_backoff.py, test_redact.py

`test_buffer_invariants.py` extended with drop-newest +
critical-event preservation cases. `test_release_polish.py`
updated to pin the 5s cap on both the sync and async jitter
helpers (post §7.2 #35 split).

Full incident write-ups in CHANGELOG.md under the same P0/S/P tags.

* fix: address ruff lint findings from CI

Three CI lint failures on `ruff check src/` — fixes only, no
behavioural changes:

- **B905** (`src/nullrun/decorators.py:162`): `zip(bound_params,
  args)` now passes `strict=False` explicitly. Pre-fix the two
  iterables can be different lengths — `bound_params` is sliced to
  `[: len(args)]` but the function may have fewer positional
  parameters than args provided (e.g. *args-style callables), in
  which case the trailing loop below handles the excess. `strict=`
  was implicit and triggered B905. Now explicit so the intent is
  documented in code.

- **I001** (`src/nullrun/instrumentation/auto.py:1146`): the late
  `import os as _os` was moved to the top-of-file import block as
  `import os` (alphabetical order: hashlib, json, logging, os,
  threading). The `_os` alias was only there to avoid shadowing —
  there is no top-level `os` in scope, so the plain name is fine.
  Call site updated to use `os.environ.get(...)`.

- **S108** (`src/nullrun/transport.py:632`): replaced the
  hardcoded `/tmp/nullrun.wal` with
  `os.path.join(tempfile.gettempdir(), "nullrun.wal")`. The
  hardcoded `/tmp` flagged S108 (insecure / non-portable temp
  path) and would have broken the SDK on Windows out of the box.
  `gettempdir()` returns the OS-appropriate temp dir
  (`/tmp` on Linux, `/var/folders/...` on macOS, `%TEMP%` on
  Windows). `NULLRUN_WAL_PATH` env override still wins, so
  containers with `readOnlyRootFilesystem: true` are unaffected.
  Added `import tempfile` to the top-of-file imports.

Verified:
  - `ruff check src/` → All checks passed!
  - `mypy src/` → Success: no issues found in 23 source files
  - `pytest` → 493 passed, 13 skipped (CI default, no `-W error`)
maltsev-dev added a commit that referenced this pull request Aug 7, 2026
* fix: P0 security/stability hardening bundle

Closes the P0/P1/P2/P3 issues from the security review (plan §10/§11.4).

Security / PCI-DSS / GDPR

- P0-1: Mask positional PII in `_enforce_sensitive_tool` by introspecting
  the wrapped function's signature and applying `SENSITIVE_ARG_KEYS` to
  positional params. Pre-fix, `charge("4111-…-1111", 50)` forwarded the
  PAN into `/execute` and the audit log.
- P0-6 / P3-3: `_safe_repr` now redacts BEFORE truncating. The pre-fix
  order truncated first, so `details={…}` past position 50 leaked
  verbatim. `_safe_repr` is now the single source of truth for the
  redact-then-truncate flow.

Cost-audit / reliability

- P0-3: Bounded chunked reads on the sync + async httpx transports
  (`MAX_RESPONSE_BYTES`, default 16 MiB, `NULLRUN_MAX_RESPONSE_BYTES`
  env override). Above the cap, tracking is skipped and
  `_coverage_streaming_skipped` is incremented. Replaces the
  `response.read()` / `await response.aread()` unbounded buffer that
  held entire LLM streaming bodies in memory.
- P0-4: `_do_flush_locked` re-queue on CB OPEN now drops the NEWEST
  non-critical events instead of the oldest. The oldest events
  (incident start, billing-period start) are exactly what a billing
  investigator needs; losing them silently broke monthly rollups.
  Control-plane events (`state_change`, `kill_received`,
  `policy_invalidated`, `key_rotated`) are preserved unconditionally
  so the dashboard KILL switch lands even under sustained backend
  outage.

Identity

- S-8 / P2-4: `agent()` now emits `str(uuid.uuid4())` (with dashes).
  Pre-fix the format was `f"agent-{uuid.uuid4().hex}"` — 32 hex chars,
  no dashes — and backend UUID-typed columns dropped these to NULL
  on insert. User-supplied names are still preserved verbatim.
- §7.2 #16: `workflow()` context manager now resets `span_id` (not
  only `workflow_id` / `trace_id`) so nested `with span()` blocks
  don't leave the inner span_id visible inside the workflow scope.

Resource leaks

- S-9: `_active_runs` on `NullRunCallback` is now an `OrderedDict`
  capped at 4096 with FIFO eviction. Pre-fix the dict grew
  unbounded when `on_chain_end` did not fire (some LangChain
  versions short-circuit the end hook on chain-body errors).
- S-10: WebSocket reconnect loop is now capped at 10 consecutive
  failures, then falls back to HTTP-poll. Pre-fix the loop ran
  forever when the backend was permanently down, leaking the
  WS thread.

Transport

- §7.2 #6: Separate `hmac_verify_expired_total` counter so SRE can
  distinguish clock-skew (NTP drift) from forged packets. Mirrored
  in both the HTTP and WebSocket verify paths.
- §7.2 #35: `CircuitBreaker.call` now dispatches the OPEN→HALF_OPEN
  jitter through `_maybe_apply_open_jitter_sync` /
  `_maybe_apply_open_jitter_async`. Pre-fix the jitter used
  `time.sleep` before dispatching to async, which blocked the
  caller's event loop on every transition.
- P2-1: `_coverage_seen` now bumps in the httpx path (sync + async).
  Pre-fix the counter was only bumped by the `requests` transport,
  so the dashboard's coverage view was empty for the dominant
  OpenAI / Anthropic / Gemini / Mistral / Cohere traffic.
- P2-3: `is_sensitive_tool` match is case-insensitive. Pre-fix
  `"stripe.charge"` did not match `"Stripe.Charge"`, bypassing the
  sensitive gate.

Concurrency

- §7.2 #39: New `_tools_lock` guards every mutation of
  `_strict_mode_tools` / `_sensitive_tools`. Same lock guards the
  coverage-counter bump+prune sequence (§7.2 #33) so two threads
  can't both observe the dict at length 4095 and both grow it to
  4097 before either prune lands.
- §7.2 #47: New `_langchain_lock` / `_langgraph_lock` guard the
  patch sequences end-to-end. Pre-fix two threads racing through
  `auto_instrument` could both pass the early `_x_patched` check
  and double-wrap `BaseCallbackManager` / `Pregel`.
- §7.2 #33: `_COVERAGE_CAP` (4096) bounds the per-host coverage
  dicts.

Webhook delivery

- P3-2: Exponential backoff (0.5s, 1s, 2s, 4s, 8s, 16s, 30s cap)
  replaces the previous linear schedule. Linear didn't back off
  fast enough under sustained outage — each KILL/PAUSE spawned
  its own delivery thread, producing 1000+ spinning threads
  hammering the dead endpoint.

WAL crash-recovery

- P1-5b: Atomic WAL writes (tmp + `fsync` + `os.replace`), 64 MiB
  rotation with `os.replace(wal, wal.1)`, replay drains both
  `wal.1` and `wal`. New `NULLRUN_WAL_PATH` / `NULLRUN_WAL_MAX_BYTES`
  env overrides for containers with `readOnlyRootFilesystem: true`.

Tests

8 new regression test files (57 tests total):
  test_agent_id_uuid.py, test_args_pii_masked.py,
  test_streaming_oom_cap.py, test_lru_active_runs.py,
  test_reconnect_cap.py, test_coverage_seen_httpx.py,
  test_webhook_backoff.py, test_redact.py

`test_buffer_invariants.py` extended with drop-newest +
critical-event preservation cases. `test_release_polish.py`
updated to pin the 5s cap on both the sync and async jitter
helpers (post §7.2 #35 split).

Full incident write-ups in CHANGELOG.md under the same P0/S/P tags.

* fix: address ruff lint findings from CI

Three CI lint failures on `ruff check src/` — fixes only, no
behavioural changes:

- **B905** (`src/nullrun/decorators.py:162`): `zip(bound_params,
  args)` now passes `strict=False` explicitly. Pre-fix the two
  iterables can be different lengths — `bound_params` is sliced to
  `[: len(args)]` but the function may have fewer positional
  parameters than args provided (e.g. *args-style callables), in
  which case the trailing loop below handles the excess. `strict=`
  was implicit and triggered B905. Now explicit so the intent is
  documented in code.

- **I001** (`src/nullrun/instrumentation/auto.py:1146`): the late
  `import os as _os` was moved to the top-of-file import block as
  `import os` (alphabetical order: hashlib, json, logging, os,
  threading). The `_os` alias was only there to avoid shadowing —
  there is no top-level `os` in scope, so the plain name is fine.
  Call site updated to use `os.environ.get(...)`.

- **S108** (`src/nullrun/transport.py:632`): replaced the
  hardcoded `/tmp/nullrun.wal` with
  `os.path.join(tempfile.gettempdir(), "nullrun.wal")`. The
  hardcoded `/tmp` flagged S108 (insecure / non-portable temp
  path) and would have broken the SDK on Windows out of the box.
  `gettempdir()` returns the OS-appropriate temp dir
  (`/tmp` on Linux, `/var/folders/...` on macOS, `%TEMP%` on
  Windows). `NULLRUN_WAL_PATH` env override still wins, so
  containers with `readOnlyRootFilesystem: true` are unaffected.
  Added `import tempfile` to the top-of-file imports.

Verified:
  - `ruff check src/` → All checks passed!
  - `mypy src/` → Success: no issues found in 23 source files
  - `pytest` → 493 passed, 13 skipped (CI default, no `-W error`)

* chore(release): bump to 0.5.2

- Promote [Unreleased] to [0.5.2] — 2026-06-19; merge the two
  [Unreleased] sections that had drifted during Sprint 2.5 +
  Phase 0 development so release tooling scanning for the
  [Unreleased] anchor picks up the complete change set exactly
  once.
- Add PEP 561 marker (py.typed) — the package ships inline type
  annotations; the marker tells mypy / pyright / pylance to honour
  them.
- runtime.py (S-4): case-insensitive state compare in
  check_control_plane. Defensive against any backend casing drift
  beyond the current PascalCase (handlers.rs:9258). Pinned by
  tests/test_state_compare_case_insensitive.py (10 cases covering
  PascalCase / UPPERCASE / lowercase / mixed-case).

Working-notes file docs/integration-baseline-2026-06-19.md is
deliberately left untracked, matching the analyze.md pattern from
d74712e.
maltsev-dev added a commit that referenced this pull request Aug 7, 2026
* fix: P0 security/stability hardening bundle

Closes the P0/P1/P2/P3 issues from the security review (plan §10/§11.4).

Security / PCI-DSS / GDPR

- P0-1: Mask positional PII in `_enforce_sensitive_tool` by introspecting
  the wrapped function's signature and applying `SENSITIVE_ARG_KEYS` to
  positional params. Pre-fix, `charge("4111-…-1111", 50)` forwarded the
  PAN into `/execute` and the audit log.
- P0-6 / P3-3: `_safe_repr` now redacts BEFORE truncating. The pre-fix
  order truncated first, so `details={…}` past position 50 leaked
  verbatim. `_safe_repr` is now the single source of truth for the
  redact-then-truncate flow.

Cost-audit / reliability

- P0-3: Bounded chunked reads on the sync + async httpx transports
  (`MAX_RESPONSE_BYTES`, default 16 MiB, `NULLRUN_MAX_RESPONSE_BYTES`
  env override). Above the cap, tracking is skipped and
  `_coverage_streaming_skipped` is incremented. Replaces the
  `response.read()` / `await response.aread()` unbounded buffer that
  held entire LLM streaming bodies in memory.
- P0-4: `_do_flush_locked` re-queue on CB OPEN now drops the NEWEST
  non-critical events instead of the oldest. The oldest events
  (incident start, billing-period start) are exactly what a billing
  investigator needs; losing them silently broke monthly rollups.
  Control-plane events (`state_change`, `kill_received`,
  `policy_invalidated`, `key_rotated`) are preserved unconditionally
  so the dashboard KILL switch lands even under sustained backend
  outage.

Identity

- S-8 / P2-4: `agent()` now emits `str(uuid.uuid4())` (with dashes).
  Pre-fix the format was `f"agent-{uuid.uuid4().hex}"` — 32 hex chars,
  no dashes — and backend UUID-typed columns dropped these to NULL
  on insert. User-supplied names are still preserved verbatim.
- §7.2 #16: `workflow()` context manager now resets `span_id` (not
  only `workflow_id` / `trace_id`) so nested `with span()` blocks
  don't leave the inner span_id visible inside the workflow scope.

Resource leaks

- S-9: `_active_runs` on `NullRunCallback` is now an `OrderedDict`
  capped at 4096 with FIFO eviction. Pre-fix the dict grew
  unbounded when `on_chain_end` did not fire (some LangChain
  versions short-circuit the end hook on chain-body errors).
- S-10: WebSocket reconnect loop is now capped at 10 consecutive
  failures, then falls back to HTTP-poll. Pre-fix the loop ran
  forever when the backend was permanently down, leaking the
  WS thread.

Transport

- §7.2 #6: Separate `hmac_verify_expired_total` counter so SRE can
  distinguish clock-skew (NTP drift) from forged packets. Mirrored
  in both the HTTP and WebSocket verify paths.
- §7.2 #35: `CircuitBreaker.call` now dispatches the OPEN→HALF_OPEN
  jitter through `_maybe_apply_open_jitter_sync` /
  `_maybe_apply_open_jitter_async`. Pre-fix the jitter used
  `time.sleep` before dispatching to async, which blocked the
  caller's event loop on every transition.
- P2-1: `_coverage_seen` now bumps in the httpx path (sync + async).
  Pre-fix the counter was only bumped by the `requests` transport,
  so the dashboard's coverage view was empty for the dominant
  OpenAI / Anthropic / Gemini / Mistral / Cohere traffic.
- P2-3: `is_sensitive_tool` match is case-insensitive. Pre-fix
  `"stripe.charge"` did not match `"Stripe.Charge"`, bypassing the
  sensitive gate.

Concurrency

- §7.2 #39: New `_tools_lock` guards every mutation of
  `_strict_mode_tools` / `_sensitive_tools`. Same lock guards the
  coverage-counter bump+prune sequence (§7.2 #33) so two threads
  can't both observe the dict at length 4095 and both grow it to
  4097 before either prune lands.
- §7.2 #47: New `_langchain_lock` / `_langgraph_lock` guard the
  patch sequences end-to-end. Pre-fix two threads racing through
  `auto_instrument` could both pass the early `_x_patched` check
  and double-wrap `BaseCallbackManager` / `Pregel`.
- §7.2 #33: `_COVERAGE_CAP` (4096) bounds the per-host coverage
  dicts.

Webhook delivery

- P3-2: Exponential backoff (0.5s, 1s, 2s, 4s, 8s, 16s, 30s cap)
  replaces the previous linear schedule. Linear didn't back off
  fast enough under sustained outage — each KILL/PAUSE spawned
  its own delivery thread, producing 1000+ spinning threads
  hammering the dead endpoint.

WAL crash-recovery

- P1-5b: Atomic WAL writes (tmp + `fsync` + `os.replace`), 64 MiB
  rotation with `os.replace(wal, wal.1)`, replay drains both
  `wal.1` and `wal`. New `NULLRUN_WAL_PATH` / `NULLRUN_WAL_MAX_BYTES`
  env overrides for containers with `readOnlyRootFilesystem: true`.

Tests

8 new regression test files (57 tests total):
  test_agent_id_uuid.py, test_args_pii_masked.py,
  test_streaming_oom_cap.py, test_lru_active_runs.py,
  test_reconnect_cap.py, test_coverage_seen_httpx.py,
  test_webhook_backoff.py, test_redact.py

`test_buffer_invariants.py` extended with drop-newest +
critical-event preservation cases. `test_release_polish.py`
updated to pin the 5s cap on both the sync and async jitter
helpers (post §7.2 #35 split).

Full incident write-ups in CHANGELOG.md under the same P0/S/P tags.

* fix: address ruff lint findings from CI

Three CI lint failures on `ruff check src/` — fixes only, no
behavioural changes:

- **B905** (`src/nullrun/decorators.py:162`): `zip(bound_params,
  args)` now passes `strict=False` explicitly. Pre-fix the two
  iterables can be different lengths — `bound_params` is sliced to
  `[: len(args)]` but the function may have fewer positional
  parameters than args provided (e.g. *args-style callables), in
  which case the trailing loop below handles the excess. `strict=`
  was implicit and triggered B905. Now explicit so the intent is
  documented in code.

- **I001** (`src/nullrun/instrumentation/auto.py:1146`): the late
  `import os as _os` was moved to the top-of-file import block as
  `import os` (alphabetical order: hashlib, json, logging, os,
  threading). The `_os` alias was only there to avoid shadowing —
  there is no top-level `os` in scope, so the plain name is fine.
  Call site updated to use `os.environ.get(...)`.

- **S108** (`src/nullrun/transport.py:632`): replaced the
  hardcoded `/tmp/nullrun.wal` with
  `os.path.join(tempfile.gettempdir(), "nullrun.wal")`. The
  hardcoded `/tmp` flagged S108 (insecure / non-portable temp
  path) and would have broken the SDK on Windows out of the box.
  `gettempdir()` returns the OS-appropriate temp dir
  (`/tmp` on Linux, `/var/folders/...` on macOS, `%TEMP%` on
  Windows). `NULLRUN_WAL_PATH` env override still wins, so
  containers with `readOnlyRootFilesystem: true` are unaffected.
  Added `import tempfile` to the top-of-file imports.

Verified:
  - `ruff check src/` → All checks passed!
  - `mypy src/` → Success: no issues found in 23 source files
  - `pytest` → 493 passed, 13 skipped (CI default, no `-W error`)

* chore(release): bump to 0.5.2

- Promote [Unreleased] to [0.5.2] — 2026-06-19; merge the two
  [Unreleased] sections that had drifted during Sprint 2.5 +
  Phase 0 development so release tooling scanning for the
  [Unreleased] anchor picks up the complete change set exactly
  once.
- Add PEP 561 marker (py.typed) — the package ships inline type
  annotations; the marker tells mypy / pyright / pylance to honour
  them.
- runtime.py (S-4): case-insensitive state compare in
  check_control_plane. Defensive against any backend casing drift
  beyond the current PascalCase (handlers.rs:9258). Pinned by
  tests/test_state_compare_case_insensitive.py (10 cases covering
  PascalCase / UPPERCASE / lowercase / mixed-case).

Working-notes file docs/integration-baseline-2026-06-19.md is
deliberately left untracked, matching the analyze.md pattern from
d74712e.

* test: bump coverage 70.92% → 84.52% with branch coverage

Lifts the SDK's Codecov score from 70.92 % to 84.52 % (+13.6 pp) by
adding 347 new tests across 10 files that exercise previously-untested
branches in the auto-instrumentation patches, runtime gates, transport
fallback modes, circuit breaker Redis path, and the @Protect decorator
fail-CLOSED contract.

pyproject.toml
  - Enable branch coverage so error / fallback paths count.
  - Raise fail_under from 70 → 82 (enforced in CI via `coverage run -m
    pytest && coverage report`).
  - Add precision=2 and skip_empty=true to keep the report readable.

New tests (all 817 pass locally, all 4 CI jobs green):

  tests/test_autogen_patch.py          — 13 tests
  tests/test_crewai_patch.py           — 15 tests
  tests/test_llama_index_patch.py      — 13 tests
  tests/test_langgraph_callback.py     — 38 tests
  tests/test_auto_requests.py          — 24 tests
  tests/test_runtime_branches.py       — 43 tests
  tests/test_transport_branches.py     — 44 tests
  tests/test_circuit_breaker_branches.py — 31 tests
  tests/test_protect_branches.py       — 43 tests
  tests/test_actions_context_init.py   — 50 tests

Per-file coverage deltas:

  instrumentation/autogen.py        21.33 → 93.41 %
  instrumentation/crewai.py         22.97 → 90.82 %
  instrumentation/llama_index.py    28.30 → 100.00 %
  instrumentation/langgraph.py      23.75 → 93.69 %
  instrumentation/auto_requests.py  33.72 → 99.09 %
  breaker/circuit_breaker.py        59.76 → 90.21 %
  transport.py                     82.57 → 84.79 %
  transport_websocket.py           68.70 → 64.10 % (msg-type branches
                                                  still need live ws
                                                  round-trip tests)
  decorators.py                    83.33 → 95.49 %
  runtime.py                       80.14 → 83.24 %
  context.py                       82.76 → 100.00 %
  actions.py                       92.12 → 96.89 %
  breaker/exceptions.py             98.51 → 97.26 %

All 4 CI jobs pass locally (pytest, ruff check, mypy, coverage).

Working-notes file docs/integration-baseline-2026-06-19.md is
deliberately left untracked, matching the analyze.md pattern from
d74712e.
maltsev-dev added a commit that referenced this pull request Aug 7, 2026
…erage reporter (#26)

* fix: P0 security/stability hardening bundle

Closes the P0/P1/P2/P3 issues from the security review (plan §10/§11.4).

Security / PCI-DSS / GDPR

- P0-1: Mask positional PII in `_enforce_sensitive_tool` by introspecting
  the wrapped function's signature and applying `SENSITIVE_ARG_KEYS` to
  positional params. Pre-fix, `charge("4111-…-1111", 50)` forwarded the
  PAN into `/execute` and the audit log.
- P0-6 / P3-3: `_safe_repr` now redacts BEFORE truncating. The pre-fix
  order truncated first, so `details={…}` past position 50 leaked
  verbatim. `_safe_repr` is now the single source of truth for the
  redact-then-truncate flow.

Cost-audit / reliability

- P0-3: Bounded chunked reads on the sync + async httpx transports
  (`MAX_RESPONSE_BYTES`, default 16 MiB, `NULLRUN_MAX_RESPONSE_BYTES`
  env override). Above the cap, tracking is skipped and
  `_coverage_streaming_skipped` is incremented. Replaces the
  `response.read()` / `await response.aread()` unbounded buffer that
  held entire LLM streaming bodies in memory.
- P0-4: `_do_flush_locked` re-queue on CB OPEN now drops the NEWEST
  non-critical events instead of the oldest. The oldest events
  (incident start, billing-period start) are exactly what a billing
  investigator needs; losing them silently broke monthly rollups.
  Control-plane events (`state_change`, `kill_received`,
  `policy_invalidated`, `key_rotated`) are preserved unconditionally
  so the dashboard KILL switch lands even under sustained backend
  outage.

Identity

- S-8 / P2-4: `agent()` now emits `str(uuid.uuid4())` (with dashes).
  Pre-fix the format was `f"agent-{uuid.uuid4().hex}"` — 32 hex chars,
  no dashes — and backend UUID-typed columns dropped these to NULL
  on insert. User-supplied names are still preserved verbatim.
- §7.2 #16: `workflow()` context manager now resets `span_id` (not
  only `workflow_id` / `trace_id`) so nested `with span()` blocks
  don't leave the inner span_id visible inside the workflow scope.

Resource leaks

- S-9: `_active_runs` on `NullRunCallback` is now an `OrderedDict`
  capped at 4096 with FIFO eviction. Pre-fix the dict grew
  unbounded when `on_chain_end` did not fire (some LangChain
  versions short-circuit the end hook on chain-body errors).
- S-10: WebSocket reconnect loop is now capped at 10 consecutive
  failures, then falls back to HTTP-poll. Pre-fix the loop ran
  forever when the backend was permanently down, leaking the
  WS thread.

Transport

- §7.2 #6: Separate `hmac_verify_expired_total` counter so SRE can
  distinguish clock-skew (NTP drift) from forged packets. Mirrored
  in both the HTTP and WebSocket verify paths.
- §7.2 #35: `CircuitBreaker.call` now dispatches the OPEN→HALF_OPEN
  jitter through `_maybe_apply_open_jitter_sync` /
  `_maybe_apply_open_jitter_async`. Pre-fix the jitter used
  `time.sleep` before dispatching to async, which blocked the
  caller's event loop on every transition.
- P2-1: `_coverage_seen` now bumps in the httpx path (sync + async).
  Pre-fix the counter was only bumped by the `requests` transport,
  so the dashboard's coverage view was empty for the dominant
  OpenAI / Anthropic / Gemini / Mistral / Cohere traffic.
- P2-3: `is_sensitive_tool` match is case-insensitive. Pre-fix
  `"stripe.charge"` did not match `"Stripe.Charge"`, bypassing the
  sensitive gate.

Concurrency

- §7.2 #39: New `_tools_lock` guards every mutation of
  `_strict_mode_tools` / `_sensitive_tools`. Same lock guards the
  coverage-counter bump+prune sequence (§7.2 #33) so two threads
  can't both observe the dict at length 4095 and both grow it to
  4097 before either prune lands.
- §7.2 #47: New `_langchain_lock` / `_langgraph_lock` guard the
  patch sequences end-to-end. Pre-fix two threads racing through
  `auto_instrument` could both pass the early `_x_patched` check
  and double-wrap `BaseCallbackManager` / `Pregel`.
- §7.2 #33: `_COVERAGE_CAP` (4096) bounds the per-host coverage
  dicts.

Webhook delivery

- P3-2: Exponential backoff (0.5s, 1s, 2s, 4s, 8s, 16s, 30s cap)
  replaces the previous linear schedule. Linear didn't back off
  fast enough under sustained outage — each KILL/PAUSE spawned
  its own delivery thread, producing 1000+ spinning threads
  hammering the dead endpoint.

WAL crash-recovery

- P1-5b: Atomic WAL writes (tmp + `fsync` + `os.replace`), 64 MiB
  rotation with `os.replace(wal, wal.1)`, replay drains both
  `wal.1` and `wal`. New `NULLRUN_WAL_PATH` / `NULLRUN_WAL_MAX_BYTES`
  env overrides for containers with `readOnlyRootFilesystem: true`.

Tests

8 new regression test files (57 tests total):
  test_agent_id_uuid.py, test_args_pii_masked.py,
  test_streaming_oom_cap.py, test_lru_active_runs.py,
  test_reconnect_cap.py, test_coverage_seen_httpx.py,
  test_webhook_backoff.py, test_redact.py

`test_buffer_invariants.py` extended with drop-newest +
critical-event preservation cases. `test_release_polish.py`
updated to pin the 5s cap on both the sync and async jitter
helpers (post §7.2 #35 split).

Full incident write-ups in CHANGELOG.md under the same P0/S/P tags.

* fix: address ruff lint findings from CI

Three CI lint failures on `ruff check src/` — fixes only, no
behavioural changes:

- **B905** (`src/nullrun/decorators.py:162`): `zip(bound_params,
  args)` now passes `strict=False` explicitly. Pre-fix the two
  iterables can be different lengths — `bound_params` is sliced to
  `[: len(args)]` but the function may have fewer positional
  parameters than args provided (e.g. *args-style callables), in
  which case the trailing loop below handles the excess. `strict=`
  was implicit and triggered B905. Now explicit so the intent is
  documented in code.

- **I001** (`src/nullrun/instrumentation/auto.py:1146`): the late
  `import os as _os` was moved to the top-of-file import block as
  `import os` (alphabetical order: hashlib, json, logging, os,
  threading). The `_os` alias was only there to avoid shadowing —
  there is no top-level `os` in scope, so the plain name is fine.
  Call site updated to use `os.environ.get(...)`.

- **S108** (`src/nullrun/transport.py:632`): replaced the
  hardcoded `/tmp/nullrun.wal` with
  `os.path.join(tempfile.gettempdir(), "nullrun.wal")`. The
  hardcoded `/tmp` flagged S108 (insecure / non-portable temp
  path) and would have broken the SDK on Windows out of the box.
  `gettempdir()` returns the OS-appropriate temp dir
  (`/tmp` on Linux, `/var/folders/...` on macOS, `%TEMP%` on
  Windows). `NULLRUN_WAL_PATH` env override still wins, so
  containers with `readOnlyRootFilesystem: true` are unaffected.
  Added `import tempfile` to the top-of-file imports.

Verified:
  - `ruff check src/` → All checks passed!
  - `mypy src/` → Success: no issues found in 23 source files
  - `pytest` → 493 passed, 13 skipped (CI default, no `-W error`)

* chore(release): bump to 0.5.2

- Promote [Unreleased] to [0.5.2] — 2026-06-19; merge the two
  [Unreleased] sections that had drifted during Sprint 2.5 +
  Phase 0 development so release tooling scanning for the
  [Unreleased] anchor picks up the complete change set exactly
  once.
- Add PEP 561 marker (py.typed) — the package ships inline type
  annotations; the marker tells mypy / pyright / pylance to honour
  them.
- runtime.py (S-4): case-insensitive state compare in
  check_control_plane. Defensive against any backend casing drift
  beyond the current PascalCase (handlers.rs:9258). Pinned by
  tests/test_state_compare_case_insensitive.py (10 cases covering
  PascalCase / UPPERCASE / lowercase / mixed-case).

Working-notes file docs/integration-baseline-2026-06-19.md is
deliberately left untracked, matching the analyze.md pattern from
d74712e.

* test: bump coverage 70.92% → 84.52% with branch coverage

Lifts the SDK's Codecov score from 70.92 % to 84.52 % (+13.6 pp) by
adding 347 new tests across 10 files that exercise previously-untested
branches in the auto-instrumentation patches, runtime gates, transport
fallback modes, circuit breaker Redis path, and the @Protect decorator
fail-CLOSED contract.

pyproject.toml
  - Enable branch coverage so error / fallback paths count.
  - Raise fail_under from 70 → 82 (enforced in CI via `coverage run -m
    pytest && coverage report`).
  - Add precision=2 and skip_empty=true to keep the report readable.

New tests (all 817 pass locally, all 4 CI jobs green):

  tests/test_autogen_patch.py          — 13 tests
  tests/test_crewai_patch.py           — 15 tests
  tests/test_llama_index_patch.py      — 13 tests
  tests/test_langgraph_callback.py     — 38 tests
  tests/test_auto_requests.py          — 24 tests
  tests/test_runtime_branches.py       — 43 tests
  tests/test_transport_branches.py     — 44 tests
  tests/test_circuit_breaker_branches.py — 31 tests
  tests/test_protect_branches.py       — 43 tests
  tests/test_actions_context_init.py   — 50 tests

Per-file coverage deltas:

  instrumentation/autogen.py        21.33 → 93.41 %
  instrumentation/crewai.py         22.97 → 90.82 %
  instrumentation/llama_index.py    28.30 → 100.00 %
  instrumentation/langgraph.py      23.75 → 93.69 %
  instrumentation/auto_requests.py  33.72 → 99.09 %
  breaker/circuit_breaker.py        59.76 → 90.21 %
  transport.py                     82.57 → 84.79 %
  transport_websocket.py           68.70 → 64.10 % (msg-type branches
                                                  still need live ws
                                                  round-trip tests)
  decorators.py                    83.33 → 95.49 %
  runtime.py                       80.14 → 83.24 %
  context.py                       82.76 → 100.00 %
  actions.py                       92.12 → 96.89 %
  breaker/exceptions.py             98.51 → 97.26 %

All 4 CI jobs pass locally (pytest, ruff check, mypy, coverage).

Working-notes file docs/integration-baseline-2026-06-19.md is
deliberately left untracked, matching the analyze.md pattern from
d74712e.

* feat(security): make @sensitive registration fail-CLOSED (ADR-008)

Sensitive-tool registration is part of the security boundary. The
old behaviour caught any exception from _get_or_create_runtime(),
logged it at DEBUG, and returned the original function unchanged —
which meant the wrapped body would later execute without ever being
added to the runtime's sensitive-tool set, completely bypassing the
pre-execution gate under partial initialization (e.g. transient
NullRunAuthenticationError on import).

Replace the silent logger.debug(...) with raise RuntimeError(...,
chained from the original exception. The decorator is the registration
point, not the call site, so raising at decoration time is the correct
signal: the import / module-load fails loudly, the body never gets a
chance to run untracked, and the caller can still inspect the root
cause via __cause__.

The two pre-existing tests pinned the old (silent / wrong-type) contract;
update them to assert the new RuntimeError wrapping:
  - test_sensitive_raises_on_missing_api_key now expects RuntimeError
    whose __cause__ is the original NullRunAuthenticationError.
  - test_sensitive_runtime_init_failure_is_silent is renamed to
    ..._raises and asserts the same __cause__ chaining when a
    _get_or_create_runtime mock raises.

* fix(transport): retry /track/batch on 5xx and align auth-verify path (P0 #2, P0 #5)

P0 #2 — _send_batch_with_retry_info used to do a single
self._client.post(...) + raise_for_status(). A transient backend 5xx
raised out of the flush path; the in-memory buffer was cleared at the
call site and every event in the batch was permanently lost. Wrap the
post() in _retry_with_backoff (max 3 attempts, exponential backoff +
jitter, capped at 10s) so a single 500 no longer drops the whole batch.
429 is retried (helper honors Retry-After when present); other 4xx
errors are returned as-is — those are real client bugs and must not
be retried (e.g. a 401 just wastes the user's budget).

P0 #5 — contract drift: this file's auth-verify call site used
/auth/verify, while the corresponding call in runtime.py:599 already
used /api/v1/auth/verify. Align the rotation call site to /api/v1/auth/verify
so the contract-drift-guard CI catches any future divergence.

Update tests/test_transport.py::test_retry_on_500 to assert the new
contract (third attempt succeeds → call_count == 3, event id in
accepted_event_ids) instead of expecting an immediate exception.
Add tests/test_track_batch_retry.py with full regression coverage:
single 5xx → success, three consecutive 5xx → BreakerTransportError,
429 with Retry-After → honored before next attempt.

* feat(runtime): emit background coverage_report every 60s

The SDK has tracked per-host seen / tracked / streaming_skipped counters
since 0.4.x (bump_coverage_counter, get_coverage_stats), but there was
no path to ship them to the backend — the counters only ever existed
in process memory. This commit adds a daemon thread that emits a
coverage_report track event every 60 seconds so the backend can build
the per-host coverage dashboard.

* NullRunRuntime.track_coverage() — returns a track-result dict when
  there is something to report, or None on cold start (no counters
  bumped yet) so the backend doesn't get an empty row per minute.
* start_coverage_reporter() / stop_coverage_reporter() — idempotent
  lifecycle, daemon thread, sleeps in 0.5s slices for responsive
  shutdown, emits once on entry so short-lived processes (CI, batch
  jobs) still leave a row.
* nullrun.init() wires start_coverage_reporter() in; the reporter is
  a no-op while the process is still cold, so re-init is safe.

New tests/test_coverage_report.py pins the contract: cold start → None,
post-traffic → track-result dict with type=coverage_report and the three
counter dicts, start is idempotent, stop joins cleanly.

* chore(breaker): add __main__ shim so 'python -m nullrun.breaker' exits cleanly

Historically the SDK shipped a 'python -m nullrun.breaker' entry point
for in-container health probes and ad-hoc debugging. The nullrun.breaker
subpackage is the circuit-breaker + policy-exceptions surface — it is
not a runnable command. Without this shim, containerized deployments
that scripted 'python -m nullrun.breaker' as a no-op smoke check would
fail with 'No module named nullrun.breaker.__main__'.

This module makes that invocation exit cleanly (return 0) and print a
short pointer to nullrun-doctor (nullrun.toolbox.diagnostics) for
real runtime checks.

* chore: gitignore audit.md (project-local working notes, sibling of analyze.md)

* test: re-align @sensitive test with fail-CLOSED contract after master merge

The auto-merge of master into this branch (commit 7875210) resolved
tests/test_protect_branches.py by taking master's side of the conflict,
leaving the old test_sensitive_runtime_init_failure_is_silent in place.
That test asserts @sensitive does NOT raise — but the production
change in commit 58263a1 (this branch) makes @sensitive raise
RuntimeError (fail-CLOSED, ADR-008). Result: CI ran the old assertion
against the new production code and failed.

Restore the renamed and re-asserted version of the test from commit
58263a1 — test_sensitive_runtime_init_failure_raises — so the test
asserts the new contract: RuntimeError is raised and __cause__ chains
the original exception.

runtime.py was resolved correctly by the auto-merge (both sides kept:
the new track_coverage / start_coverage_reporter / stop_coverage_reporter
/ _coverage_reporter_loop methods AND the existing bump_coverage_counter
are all present), so no changes there.
maltsev-dev added a commit that referenced this pull request Aug 7, 2026
* fix: P0 security/stability hardening bundle

Closes the P0/P1/P2/P3 issues from the security review (plan §10/§11.4).

Security / PCI-DSS / GDPR

- P0-1: Mask positional PII in `_enforce_sensitive_tool` by introspecting
  the wrapped function's signature and applying `SENSITIVE_ARG_KEYS` to
  positional params. Pre-fix, `charge("4111-…-1111", 50)` forwarded the
  PAN into `/execute` and the audit log.
- P0-6 / P3-3: `_safe_repr` now redacts BEFORE truncating. The pre-fix
  order truncated first, so `details={…}` past position 50 leaked
  verbatim. `_safe_repr` is now the single source of truth for the
  redact-then-truncate flow.

Cost-audit / reliability

- P0-3: Bounded chunked reads on the sync + async httpx transports
  (`MAX_RESPONSE_BYTES`, default 16 MiB, `NULLRUN_MAX_RESPONSE_BYTES`
  env override). Above the cap, tracking is skipped and
  `_coverage_streaming_skipped` is incremented. Replaces the
  `response.read()` / `await response.aread()` unbounded buffer that
  held entire LLM streaming bodies in memory.
- P0-4: `_do_flush_locked` re-queue on CB OPEN now drops the NEWEST
  non-critical events instead of the oldest. The oldest events
  (incident start, billing-period start) are exactly what a billing
  investigator needs; losing them silently broke monthly rollups.
  Control-plane events (`state_change`, `kill_received`,
  `policy_invalidated`, `key_rotated`) are preserved unconditionally
  so the dashboard KILL switch lands even under sustained backend
  outage.

Identity

- S-8 / P2-4: `agent()` now emits `str(uuid.uuid4())` (with dashes).
  Pre-fix the format was `f"agent-{uuid.uuid4().hex}"` — 32 hex chars,
  no dashes — and backend UUID-typed columns dropped these to NULL
  on insert. User-supplied names are still preserved verbatim.
- §7.2 #16: `workflow()` context manager now resets `span_id` (not
  only `workflow_id` / `trace_id`) so nested `with span()` blocks
  don't leave the inner span_id visible inside the workflow scope.

Resource leaks

- S-9: `_active_runs` on `NullRunCallback` is now an `OrderedDict`
  capped at 4096 with FIFO eviction. Pre-fix the dict grew
  unbounded when `on_chain_end` did not fire (some LangChain
  versions short-circuit the end hook on chain-body errors).
- S-10: WebSocket reconnect loop is now capped at 10 consecutive
  failures, then falls back to HTTP-poll. Pre-fix the loop ran
  forever when the backend was permanently down, leaking the
  WS thread.

Transport

- §7.2 #6: Separate `hmac_verify_expired_total` counter so SRE can
  distinguish clock-skew (NTP drift) from forged packets. Mirrored
  in both the HTTP and WebSocket verify paths.
- §7.2 #35: `CircuitBreaker.call` now dispatches the OPEN→HALF_OPEN
  jitter through `_maybe_apply_open_jitter_sync` /
  `_maybe_apply_open_jitter_async`. Pre-fix the jitter used
  `time.sleep` before dispatching to async, which blocked the
  caller's event loop on every transition.
- P2-1: `_coverage_seen` now bumps in the httpx path (sync + async).
  Pre-fix the counter was only bumped by the `requests` transport,
  so the dashboard's coverage view was empty for the dominant
  OpenAI / Anthropic / Gemini / Mistral / Cohere traffic.
- P2-3: `is_sensitive_tool` match is case-insensitive. Pre-fix
  `"stripe.charge"` did not match `"Stripe.Charge"`, bypassing the
  sensitive gate.

Concurrency

- §7.2 #39: New `_tools_lock` guards every mutation of
  `_strict_mode_tools` / `_sensitive_tools`. Same lock guards the
  coverage-counter bump+prune sequence (§7.2 #33) so two threads
  can't both observe the dict at length 4095 and both grow it to
  4097 before either prune lands.
- §7.2 #47: New `_langchain_lock` / `_langgraph_lock` guard the
  patch sequences end-to-end. Pre-fix two threads racing through
  `auto_instrument` could both pass the early `_x_patched` check
  and double-wrap `BaseCallbackManager` / `Pregel`.
- §7.2 #33: `_COVERAGE_CAP` (4096) bounds the per-host coverage
  dicts.

Webhook delivery

- P3-2: Exponential backoff (0.5s, 1s, 2s, 4s, 8s, 16s, 30s cap)
  replaces the previous linear schedule. Linear didn't back off
  fast enough under sustained outage — each KILL/PAUSE spawned
  its own delivery thread, producing 1000+ spinning threads
  hammering the dead endpoint.

WAL crash-recovery

- P1-5b: Atomic WAL writes (tmp + `fsync` + `os.replace`), 64 MiB
  rotation with `os.replace(wal, wal.1)`, replay drains both
  `wal.1` and `wal`. New `NULLRUN_WAL_PATH` / `NULLRUN_WAL_MAX_BYTES`
  env overrides for containers with `readOnlyRootFilesystem: true`.

Tests

8 new regression test files (57 tests total):
  test_agent_id_uuid.py, test_args_pii_masked.py,
  test_streaming_oom_cap.py, test_lru_active_runs.py,
  test_reconnect_cap.py, test_coverage_seen_httpx.py,
  test_webhook_backoff.py, test_redact.py

`test_buffer_invariants.py` extended with drop-newest +
critical-event preservation cases. `test_release_polish.py`
updated to pin the 5s cap on both the sync and async jitter
helpers (post §7.2 #35 split).

Full incident write-ups in CHANGELOG.md under the same P0/S/P tags.

* fix: address ruff lint findings from CI

Three CI lint failures on `ruff check src/` — fixes only, no
behavioural changes:

- **B905** (`src/nullrun/decorators.py:162`): `zip(bound_params,
  args)` now passes `strict=False` explicitly. Pre-fix the two
  iterables can be different lengths — `bound_params` is sliced to
  `[: len(args)]` but the function may have fewer positional
  parameters than args provided (e.g. *args-style callables), in
  which case the trailing loop below handles the excess. `strict=`
  was implicit and triggered B905. Now explicit so the intent is
  documented in code.

- **I001** (`src/nullrun/instrumentation/auto.py:1146`): the late
  `import os as _os` was moved to the top-of-file import block as
  `import os` (alphabetical order: hashlib, json, logging, os,
  threading). The `_os` alias was only there to avoid shadowing —
  there is no top-level `os` in scope, so the plain name is fine.
  Call site updated to use `os.environ.get(...)`.

- **S108** (`src/nullrun/transport.py:632`): replaced the
  hardcoded `/tmp/nullrun.wal` with
  `os.path.join(tempfile.gettempdir(), "nullrun.wal")`. The
  hardcoded `/tmp` flagged S108 (insecure / non-portable temp
  path) and would have broken the SDK on Windows out of the box.
  `gettempdir()` returns the OS-appropriate temp dir
  (`/tmp` on Linux, `/var/folders/...` on macOS, `%TEMP%` on
  Windows). `NULLRUN_WAL_PATH` env override still wins, so
  containers with `readOnlyRootFilesystem: true` are unaffected.
  Added `import tempfile` to the top-of-file imports.

Verified:
  - `ruff check src/` → All checks passed!
  - `mypy src/` → Success: no issues found in 23 source files
  - `pytest` → 493 passed, 13 skipped (CI default, no `-W error`)

* chore(release): bump to 0.5.2

- Promote [Unreleased] to [0.5.2] — 2026-06-19; merge the two
  [Unreleased] sections that had drifted during Sprint 2.5 +
  Phase 0 development so release tooling scanning for the
  [Unreleased] anchor picks up the complete change set exactly
  once.
- Add PEP 561 marker (py.typed) — the package ships inline type
  annotations; the marker tells mypy / pyright / pylance to honour
  them.
- runtime.py (S-4): case-insensitive state compare in
  check_control_plane. Defensive against any backend casing drift
  beyond the current PascalCase (handlers.rs:9258). Pinned by
  tests/test_state_compare_case_insensitive.py (10 cases covering
  PascalCase / UPPERCASE / lowercase / mixed-case).

Working-notes file docs/integration-baseline-2026-06-19.md is
deliberately left untracked, matching the analyze.md pattern from
d74712e.

* test: bump coverage 70.92% → 84.52% with branch coverage

Lifts the SDK's Codecov score from 70.92 % to 84.52 % (+13.6 pp) by
adding 347 new tests across 10 files that exercise previously-untested
branches in the auto-instrumentation patches, runtime gates, transport
fallback modes, circuit breaker Redis path, and the @Protect decorator
fail-CLOSED contract.

pyproject.toml
  - Enable branch coverage so error / fallback paths count.
  - Raise fail_under from 70 → 82 (enforced in CI via `coverage run -m
    pytest && coverage report`).
  - Add precision=2 and skip_empty=true to keep the report readable.

New tests (all 817 pass locally, all 4 CI jobs green):

  tests/test_autogen_patch.py          — 13 tests
  tests/test_crewai_patch.py           — 15 tests
  tests/test_llama_index_patch.py      — 13 tests
  tests/test_langgraph_callback.py     — 38 tests
  tests/test_auto_requests.py          — 24 tests
  tests/test_runtime_branches.py       — 43 tests
  tests/test_transport_branches.py     — 44 tests
  tests/test_circuit_breaker_branches.py — 31 tests
  tests/test_protect_branches.py       — 43 tests
  tests/test_actions_context_init.py   — 50 tests

Per-file coverage deltas:

  instrumentation/autogen.py        21.33 → 93.41 %
  instrumentation/crewai.py         22.97 → 90.82 %
  instrumentation/llama_index.py    28.30 → 100.00 %
  instrumentation/langgraph.py      23.75 → 93.69 %
  instrumentation/auto_requests.py  33.72 → 99.09 %
  breaker/circuit_breaker.py        59.76 → 90.21 %
  transport.py                     82.57 → 84.79 %
  transport_websocket.py           68.70 → 64.10 % (msg-type branches
                                                  still need live ws
                                                  round-trip tests)
  decorators.py                    83.33 → 95.49 %
  runtime.py                       80.14 → 83.24 %
  context.py                       82.76 → 100.00 %
  actions.py                       92.12 → 96.89 %
  breaker/exceptions.py             98.51 → 97.26 %

All 4 CI jobs pass locally (pytest, ruff check, mypy, coverage).

Working-notes file docs/integration-baseline-2026-06-19.md is
deliberately left untracked, matching the analyze.md pattern from
d74712e.

* feat(security): make @sensitive registration fail-CLOSED (ADR-008)

Sensitive-tool registration is part of the security boundary. The
old behaviour caught any exception from _get_or_create_runtime(),
logged it at DEBUG, and returned the original function unchanged —
which meant the wrapped body would later execute without ever being
added to the runtime's sensitive-tool set, completely bypassing the
pre-execution gate under partial initialization (e.g. transient
NullRunAuthenticationError on import).

Replace the silent logger.debug(...) with raise RuntimeError(...,
chained from the original exception. The decorator is the registration
point, not the call site, so raising at decoration time is the correct
signal: the import / module-load fails loudly, the body never gets a
chance to run untracked, and the caller can still inspect the root
cause via __cause__.

The two pre-existing tests pinned the old (silent / wrong-type) contract;
update them to assert the new RuntimeError wrapping:
  - test_sensitive_raises_on_missing_api_key now expects RuntimeError
    whose __cause__ is the original NullRunAuthenticationError.
  - test_sensitive_runtime_init_failure_is_silent is renamed to
    ..._raises and asserts the same __cause__ chaining when a
    _get_or_create_runtime mock raises.

* fix(transport): retry /track/batch on 5xx and align auth-verify path (P0 #2, P0 #5)

P0 #2 — _send_batch_with_retry_info used to do a single
self._client.post(...) + raise_for_status(). A transient backend 5xx
raised out of the flush path; the in-memory buffer was cleared at the
call site and every event in the batch was permanently lost. Wrap the
post() in _retry_with_backoff (max 3 attempts, exponential backoff +
jitter, capped at 10s) so a single 500 no longer drops the whole batch.
429 is retried (helper honors Retry-After when present); other 4xx
errors are returned as-is — those are real client bugs and must not
be retried (e.g. a 401 just wastes the user's budget).

P0 #5 — contract drift: this file's auth-verify call site used
/auth/verify, while the corresponding call in runtime.py:599 already
used /api/v1/auth/verify. Align the rotation call site to /api/v1/auth/verify
so the contract-drift-guard CI catches any future divergence.

Update tests/test_transport.py::test_retry_on_500 to assert the new
contract (third attempt succeeds → call_count == 3, event id in
accepted_event_ids) instead of expecting an immediate exception.
Add tests/test_track_batch_retry.py with full regression coverage:
single 5xx → success, three consecutive 5xx → BreakerTransportError,
429 with Retry-After → honored before next attempt.

* feat(runtime): emit background coverage_report every 60s

The SDK has tracked per-host seen / tracked / streaming_skipped counters
since 0.4.x (bump_coverage_counter, get_coverage_stats), but there was
no path to ship them to the backend — the counters only ever existed
in process memory. This commit adds a daemon thread that emits a
coverage_report track event every 60 seconds so the backend can build
the per-host coverage dashboard.

* NullRunRuntime.track_coverage() — returns a track-result dict when
  there is something to report, or None on cold start (no counters
  bumped yet) so the backend doesn't get an empty row per minute.
* start_coverage_reporter() / stop_coverage_reporter() — idempotent
  lifecycle, daemon thread, sleeps in 0.5s slices for responsive
  shutdown, emits once on entry so short-lived processes (CI, batch
  jobs) still leave a row.
* nullrun.init() wires start_coverage_reporter() in; the reporter is
  a no-op while the process is still cold, so re-init is safe.

New tests/test_coverage_report.py pins the contract: cold start → None,
post-traffic → track-result dict with type=coverage_report and the three
counter dicts, start is idempotent, stop joins cleanly.

* chore(breaker): add __main__ shim so 'python -m nullrun.breaker' exits cleanly

Historically the SDK shipped a 'python -m nullrun.breaker' entry point
for in-container health probes and ad-hoc debugging. The nullrun.breaker
subpackage is the circuit-breaker + policy-exceptions surface — it is
not a runnable command. Without this shim, containerized deployments
that scripted 'python -m nullrun.breaker' as a no-op smoke check would
fail with 'No module named nullrun.breaker.__main__'.

This module makes that invocation exit cleanly (return 0) and print a
short pointer to nullrun-doctor (nullrun.toolbox.diagnostics) for
real runtime checks.

* chore: gitignore audit.md (project-local working notes, sibling of analyze.md)

* test: re-align @sensitive test with fail-CLOSED contract after master merge

The auto-merge of master into this branch (commit 7875210) resolved
tests/test_protect_branches.py by taking master's side of the conflict,
leaving the old test_sensitive_runtime_init_failure_is_silent in place.
That test asserts @sensitive does NOT raise — but the production
change in commit 58263a1 (this branch) makes @sensitive raise
RuntimeError (fail-CLOSED, ADR-008). Result: CI ran the old assertion
against the new production code and failed.

Restore the renamed and re-asserted version of the test from commit
58263a1 — test_sensitive_runtime_init_failure_raises — so the test
asserts the new contract: RuntimeError is raised and __cause__ chains
the original exception.

runtime.py was resolved correctly by the auto-merge (both sides kept:
the new track_coverage / start_coverage_reporter / stop_coverage_reporter
/ _coverage_reporter_loop methods AND the existing bump_coverage_counter
are all present), so no changes there.
maltsev-dev added a commit that referenced this pull request Aug 7, 2026
* release: 0.7.6 — FastAPI integration + user-facing message catalog

Additive patch on top of the 0.7.0 thin-client refactor. No
breaking changes.

Added
-----

* nullrun.integrations.fastapi — one-line FastAPI integration
  that turns every NullRunDecision / NullRunInfrastructureError
  thrown by @nullrun.protect endpoints into a clean JSON
  response with the right HTTP status code. No per-endpoint
  except blocks required.

  Response shape:
    {"error_code": "NR-B004",
     "user_message": "You've reached the usage limit...",
     "category": "decision"}

  HTTP status mapping:
    * NR-B004 (budget), NR-L001 (loop), NR-R001 (rate) -> 429
      with optional Retry-After
    * NR-T001 (tool blocked), NR-X001 (generic block) -> 403
    * NR-W003 (paused) -> 503 with Retry-After
    * NR-W002 (killed) -> 503; WorkflowKilledInterrupt is a
      BaseException subclass so Starlette's
      add_exception_handler refuses it — handled via ASGI
      middleware instead (hybrid pattern, documented in
      module docstring).
    * NullRunInfrastructureError subclasses -> 503 (our side,
      not user's).

* nullrun.messages — default user-facing message catalog.
  Every NR-* error code has an English default message owned
  by NULLRUN, not customer code. Customer Support Bots hitting
  a budget cap show the same wording across every NullRun-backed
  application.
    * format_user_message(exc) — render exception as user-facing
      string
    * set_user_message(code, text) — per-process override for
      branded variants
    * get_user_message(code) — raw lookup
    * reset_overrides() — clear all overrides (for tests)

Changed
-------

* Transport._send_batch canonical JSON serialization — route the
  /track/batch body through _signed_request_body for consistent
  compact-separator serialization. HMAC itself is unaffected,
  but consistent serialization removes a special-case from the
  wire-format contract tests.

* Transport._send_batch actions response handling — backend
  renamed BatchTrackResponse.actions_taken (debug names) ->
  BatchTrackResponse.actions (ActionTaken structs). Read both
  for forward-compat; per-element try/except so one malformed
  entry doesn't abort the whole loop.

* pyproject.toml metadata — long-form description with search
  keywords, Maintainer: populated via maintainers=[...],
  expanded classifiers (Linux / Windows / macOS, Python 3.13,
  CPython, Security / AI / WWW/HTTP topics), project URL
  expander.

Tests
-----

* tests/test_messages.py (new, 282 lines) — catalog
  completeness (every NR-* code has a default message),
  override / reset behavior, render path.
* tests/test_integrations_fastapi.py (new, 289 lines) — HTTP
  status mapping per error code, response shape, ASGI
  middleware path for WorkflowKilledInterrupt, hybrid
  composition.
* tests/test_decision_split.py (new, 199 lines) — pins the
  decision / infrastructure error split.
* Updates to tests/test_runtime.py, tests/test_extractors.py
  reflecting transport canonical-JSON + actions-renamed
  changes.

Release plumbing
----------------

* pyproject.toml: version bumped 0.7.0 -> 0.7.6
* src/nullrun/__version__.py: __version__ = "0.7.6"
* CHANGELOG.md: full 0.7.6 entry covering additions,
  transport changes, metadata improvements

Tests pass locally (per session log) — pytest on Windows /
Python 3.14.2 is green.

* ci: fix PR #35 — fastapi dep + Transport._send_batch typo + coverage padding

PR #35 (release/0.7.6) failed all four CI jobs (test 3.10/3.11/3.12,
coverage, codecov/patch) on the same root cause + one latent bug
masked by it. This commit lands the fixes plus the last-mile tests
that bring coverage above the 82% threshold.

CI failure root
---------------

* tests/test_integrations_fastapi.py does from fastapi import ...
  at module top-level. CI installs only pip install -e '.[dev]',
  and fastapi was declared as an *optional* [fastapi] extra,
  NOT in [dev]. Pytest collection aborted with
  ModuleNotFoundError: No module named 'fastapi' → all 4 jobs red.
* Fix: add fastapi>=0.100,<1.0 to [dev]. Same precedent as
  langchain-core (already in [dev] for the same import-time
  contract: nullrun.instrumentation.langgraph is eager-imported
  from nullrun.decorators at collection time, so the test extras
  must cover the import chain).

Latent bug surfaced by the first fix
------------------------------------

The same PR refactored Transport._send_batch_with_retry_info to
route the /track/batch body through _signed_request_body for
canonical-JSON serialization (matching /gate and /execute). The two
sibling call sites use the module-level helper _signed_request_body
(no self.); this one used self._signed_request_body by typo.
Result: AttributeError on every batch flush, breaking 15 existing
tests across test_transport.py / test_track_batch_retry.py /
test_integration_contract.py / test_signal_safety.py. As long as
the fastapi collection error aborted pytest, this was hidden. Fixed
to _signed_request_body(...) with a docstring noting why it is
module-level and what the bug looked like.

Coverage padding (codecov/patch was failing on this too)
--------------------------------------------------------

Total coverage on the failing CI run was 81.98% — 0.02pp under the
fail-under=82 gate. After the two fixes above it would have
recovered to ~82.0% on the dot, so I added minimal tests for the
cheapest-to-cover gaps:

* tests/test_breaker_main.py (new) — covers the 5 statements in
  nullrun.breaker.__main__.main() (0% → 100%). The module
  exists so python -m nullrun.breaker exits cleanly instead of
  failing with No module named nullrun.breaker.__main__; the
  previous fix-mechanism was return 0 after a print, but no
  test was exercising it.
* tests/test_status.py — extends TestSummary with seven
  scenarios covering each conditional branch of NullRunStatus.summary()
  (organization_id, workflow_id, workflow_state != Normal,
  backend_reachable=False, ws_connected=False, recent_errors).
  status.py jumps 84.52% → 98.81%.
* tests/test_integrations_fastapi.py — four tests on
  _build_headers covering non-numeric, zero, negative, and
  resume_after (the WorkflowPausedException code path).
  integrations/fastapi.py jumps 90.22% → 94.57%.

After all three: TOTAL 81.98% → 82.46%, comfortably above the gate.

Verification
------------

* Local pytest: 997 passed, 13 skipped, 0 failed
  (Windows / Python 3.14.2, 8m47s — same env the original commit
  was validated in).
* python -m coverage report — 82.46%, no fail-under complaint.

* test: cover Phase 4.1 instrumentation — finish_reason + cache/reasoning/tools

Patch coverage on PR #35 was 62.38% against a 65% threshold (codecov
target 70% / threshold 5pp). The two biggest delta-holders against
master were auto.py (+286) and langgraph.py (+221), both dominated
by Phase 4.1 additions:

  * auto._normalize_finish_reason + _FINISH_REASON_MAP
  * auto._openai_extractor  second-tier fields (cache_read_tokens,
    cache_write_tokens, reasoning_tokens, finish_reason, tool_names)
  * auto._anthropic_extractor cache_read / cache_write
  * langgraph._safe_get_gen_message
  * langgraph._get_finish_reason (5-source fallback chain)
  * langgraph.extract_usage_from_response second-tier fields

These are pure / near-pure functions with no network or vendor SDK
calls. Coverage padding is cheap — pin the canonical wire shapes
once and the backend ingest contract gets a free live spec.

Local numbers:
  * auto.py        63.44% -> 64.01%   (file-level, +57 statements)
  * langgraph.py   78.50% -> 86.01%   (file-level, +32 statements)
  * TOTAL          82.46% -> 83.13%   (already above 82% gate)

41 tests, all green. Existing test_extractors.py and
test_langgraph_callback.py left untouched — these tests
deliberately target the Phase 4.1 fields (cache_read /
cache_write / reasoning / finish_reason / tool_names) that the
older tests didn't pin.
maltsev-dev added a commit that referenced this pull request Aug 7, 2026
…36)

* release: 0.7.6 — FastAPI integration + user-facing message catalog

Additive patch on top of the 0.7.0 thin-client refactor. No
breaking changes.

Added
-----

* nullrun.integrations.fastapi — one-line FastAPI integration
  that turns every NullRunDecision / NullRunInfrastructureError
  thrown by @nullrun.protect endpoints into a clean JSON
  response with the right HTTP status code. No per-endpoint
  except blocks required.

  Response shape:
    {"error_code": "NR-B004",
     "user_message": "You've reached the usage limit...",
     "category": "decision"}

  HTTP status mapping:
    * NR-B004 (budget), NR-L001 (loop), NR-R001 (rate) -> 429
      with optional Retry-After
    * NR-T001 (tool blocked), NR-X001 (generic block) -> 403
    * NR-W003 (paused) -> 503 with Retry-After
    * NR-W002 (killed) -> 503; WorkflowKilledInterrupt is a
      BaseException subclass so Starlette's
      add_exception_handler refuses it — handled via ASGI
      middleware instead (hybrid pattern, documented in
      module docstring).
    * NullRunInfrastructureError subclasses -> 503 (our side,
      not user's).

* nullrun.messages — default user-facing message catalog.
  Every NR-* error code has an English default message owned
  by NULLRUN, not customer code. Customer Support Bots hitting
  a budget cap show the same wording across every NullRun-backed
  application.
    * format_user_message(exc) — render exception as user-facing
      string
    * set_user_message(code, text) — per-process override for
      branded variants
    * get_user_message(code) — raw lookup
    * reset_overrides() — clear all overrides (for tests)

Changed
-------

* Transport._send_batch canonical JSON serialization — route the
  /track/batch body through _signed_request_body for consistent
  compact-separator serialization. HMAC itself is unaffected,
  but consistent serialization removes a special-case from the
  wire-format contract tests.

* Transport._send_batch actions response handling — backend
  renamed BatchTrackResponse.actions_taken (debug names) ->
  BatchTrackResponse.actions (ActionTaken structs). Read both
  for forward-compat; per-element try/except so one malformed
  entry doesn't abort the whole loop.

* pyproject.toml metadata — long-form description with search
  keywords, Maintainer: populated via maintainers=[...],
  expanded classifiers (Linux / Windows / macOS, Python 3.13,
  CPython, Security / AI / WWW/HTTP topics), project URL
  expander.

Tests
-----

* tests/test_messages.py (new, 282 lines) — catalog
  completeness (every NR-* code has a default message),
  override / reset behavior, render path.
* tests/test_integrations_fastapi.py (new, 289 lines) — HTTP
  status mapping per error code, response shape, ASGI
  middleware path for WorkflowKilledInterrupt, hybrid
  composition.
* tests/test_decision_split.py (new, 199 lines) — pins the
  decision / infrastructure error split.
* Updates to tests/test_runtime.py, tests/test_extractors.py
  reflecting transport canonical-JSON + actions-renamed
  changes.

Release plumbing
----------------

* pyproject.toml: version bumped 0.7.0 -> 0.7.6
* src/nullrun/__version__.py: __version__ = "0.7.6"
* CHANGELOG.md: full 0.7.6 entry covering additions,
  transport changes, metadata improvements

Tests pass locally (per session log) — pytest on Windows /
Python 3.14.2 is green.

* ci: fix PR #35 — fastapi dep + Transport._send_batch typo + coverage padding

PR #35 (release/0.7.6) failed all four CI jobs (test 3.10/3.11/3.12,
coverage, codecov/patch) on the same root cause + one latent bug
masked by it. This commit lands the fixes plus the last-mile tests
that bring coverage above the 82% threshold.

CI failure root
---------------

* tests/test_integrations_fastapi.py does from fastapi import ...
  at module top-level. CI installs only pip install -e '.[dev]',
  and fastapi was declared as an *optional* [fastapi] extra,
  NOT in [dev]. Pytest collection aborted with
  ModuleNotFoundError: No module named 'fastapi' → all 4 jobs red.
* Fix: add fastapi>=0.100,<1.0 to [dev]. Same precedent as
  langchain-core (already in [dev] for the same import-time
  contract: nullrun.instrumentation.langgraph is eager-imported
  from nullrun.decorators at collection time, so the test extras
  must cover the import chain).

Latent bug surfaced by the first fix
------------------------------------

The same PR refactored Transport._send_batch_with_retry_info to
route the /track/batch body through _signed_request_body for
canonical-JSON serialization (matching /gate and /execute). The two
sibling call sites use the module-level helper _signed_request_body
(no self.); this one used self._signed_request_body by typo.
Result: AttributeError on every batch flush, breaking 15 existing
tests across test_transport.py / test_track_batch_retry.py /
test_integration_contract.py / test_signal_safety.py. As long as
the fastapi collection error aborted pytest, this was hidden. Fixed
to _signed_request_body(...) with a docstring noting why it is
module-level and what the bug looked like.

Coverage padding (codecov/patch was failing on this too)
--------------------------------------------------------

Total coverage on the failing CI run was 81.98% — 0.02pp under the
fail-under=82 gate. After the two fixes above it would have
recovered to ~82.0% on the dot, so I added minimal tests for the
cheapest-to-cover gaps:

* tests/test_breaker_main.py (new) — covers the 5 statements in
  nullrun.breaker.__main__.main() (0% → 100%). The module
  exists so python -m nullrun.breaker exits cleanly instead of
  failing with No module named nullrun.breaker.__main__; the
  previous fix-mechanism was return 0 after a print, but no
  test was exercising it.
* tests/test_status.py — extends TestSummary with seven
  scenarios covering each conditional branch of NullRunStatus.summary()
  (organization_id, workflow_id, workflow_state != Normal,
  backend_reachable=False, ws_connected=False, recent_errors).
  status.py jumps 84.52% → 98.81%.
* tests/test_integrations_fastapi.py — four tests on
  _build_headers covering non-numeric, zero, negative, and
  resume_after (the WorkflowPausedException code path).
  integrations/fastapi.py jumps 90.22% → 94.57%.

After all three: TOTAL 81.98% → 82.46%, comfortably above the gate.

Verification
------------

* Local pytest: 997 passed, 13 skipped, 0 failed
  (Windows / Python 3.14.2, 8m47s — same env the original commit
  was validated in).
* python -m coverage report — 82.46%, no fail-under complaint.

* test: cover Phase 4.1 instrumentation — finish_reason + cache/reasoning/tools

Patch coverage on PR #35 was 62.38% against a 65% threshold (codecov
target 70% / threshold 5pp). The two biggest delta-holders against
master were auto.py (+286) and langgraph.py (+221), both dominated
by Phase 4.1 additions:

  * auto._normalize_finish_reason + _FINISH_REASON_MAP
  * auto._openai_extractor  second-tier fields (cache_read_tokens,
    cache_write_tokens, reasoning_tokens, finish_reason, tool_names)
  * auto._anthropic_extractor cache_read / cache_write
  * langgraph._safe_get_gen_message
  * langgraph._get_finish_reason (5-source fallback chain)
  * langgraph.extract_usage_from_response second-tier fields

These are pure / near-pure functions with no network or vendor SDK
calls. Coverage padding is cheap — pin the canonical wire shapes
once and the backend ingest contract gets a free live spec.

Local numbers:
  * auto.py        63.44% -> 64.01%   (file-level, +57 statements)
  * langgraph.py   78.50% -> 86.01%   (file-level, +32 statements)
  * TOTAL          82.46% -> 83.13%   (already above 82% gate)

41 tests, all green. Existing test_extractors.py and
test_langgraph_callback.py left untouched — these tests
deliberately target the Phase 4.1 fields (cache_read /
cache_write / reasoning / finish_reason / tool_names) that the
older tests didn't pin.

* fix(gate): forward real model + tools to /gate pre-flight (T4)

Pre-0.7.7 every SDK /gate call for any workflow with a budget was

hard-blocked because the runtime hard-coded the literal string

"budget-precheck" as the model. The backend's PolicyEvaluationGraph

treated any synthetic cost_limit rule with score > 0.8 as Block,

so the pricing lookup never landed on a real model and the rule

fired with the wrong score.

This commit:

* Adds nullrun.set_call_context(model=..., tools=[...]) plus

  get_call_model / get_call_tools helpers (and the underlying

  _call_model_var / _call_tools_var contextvars in

  nullrun.context).

* Wires the call context into check_workflow_budget: the /gate

  payload now carries the real model name (or None when unset)

  and the user-supplied tool list. tools=[] vs missing-None are

  distinguished on the wire per gate/internal.rs::check_tool_block.

* Transport.check forwards the tools key when set (it was

  silently dropped pre-fix).

* tests/conftest.py reset_runtime clears the new contextvars so

  a test's set_call_context(...) doesn't leak into the next

  test's wire payload.

* New tests/test_gate_real_path.py pins down the regression:

  default request allows a clean workflow, real block still

  honored, no policy-N residue on the wire, set_call_context

  flows into the body, no-context means no tools key, and the

  helpers are reachable from nullrun.*.

Bumps version to 0.7.7. No breaking changes - new helpers

default to None / empty so existing call sites keep working.
maltsev-dev added a commit that referenced this pull request Aug 7, 2026
* release: 0.7.6 — FastAPI integration + user-facing message catalog

Additive patch on top of the 0.7.0 thin-client refactor. No
breaking changes.

Added
-----

* nullrun.integrations.fastapi — one-line FastAPI integration
  that turns every NullRunDecision / NullRunInfrastructureError
  thrown by @nullrun.protect endpoints into a clean JSON
  response with the right HTTP status code. No per-endpoint
  except blocks required.

  Response shape:
    {"error_code": "NR-B004",
     "user_message": "You've reached the usage limit...",
     "category": "decision"}

  HTTP status mapping:
    * NR-B004 (budget), NR-L001 (loop), NR-R001 (rate) -> 429
      with optional Retry-After
    * NR-T001 (tool blocked), NR-X001 (generic block) -> 403
    * NR-W003 (paused) -> 503 with Retry-After
    * NR-W002 (killed) -> 503; WorkflowKilledInterrupt is a
      BaseException subclass so Starlette's
      add_exception_handler refuses it — handled via ASGI
      middleware instead (hybrid pattern, documented in
      module docstring).
    * NullRunInfrastructureError subclasses -> 503 (our side,
      not user's).

* nullrun.messages — default user-facing message catalog.
  Every NR-* error code has an English default message owned
  by NULLRUN, not customer code. Customer Support Bots hitting
  a budget cap show the same wording across every NullRun-backed
  application.
    * format_user_message(exc) — render exception as user-facing
      string
    * set_user_message(code, text) — per-process override for
      branded variants
    * get_user_message(code) — raw lookup
    * reset_overrides() — clear all overrides (for tests)

Changed
-------

* Transport._send_batch canonical JSON serialization — route the
  /track/batch body through _signed_request_body for consistent
  compact-separator serialization. HMAC itself is unaffected,
  but consistent serialization removes a special-case from the
  wire-format contract tests.

* Transport._send_batch actions response handling — backend
  renamed BatchTrackResponse.actions_taken (debug names) ->
  BatchTrackResponse.actions (ActionTaken structs). Read both
  for forward-compat; per-element try/except so one malformed
  entry doesn't abort the whole loop.

* pyproject.toml metadata — long-form description with search
  keywords, Maintainer: populated via maintainers=[...],
  expanded classifiers (Linux / Windows / macOS, Python 3.13,
  CPython, Security / AI / WWW/HTTP topics), project URL
  expander.

Tests
-----

* tests/test_messages.py (new, 282 lines) — catalog
  completeness (every NR-* code has a default message),
  override / reset behavior, render path.
* tests/test_integrations_fastapi.py (new, 289 lines) — HTTP
  status mapping per error code, response shape, ASGI
  middleware path for WorkflowKilledInterrupt, hybrid
  composition.
* tests/test_decision_split.py (new, 199 lines) — pins the
  decision / infrastructure error split.
* Updates to tests/test_runtime.py, tests/test_extractors.py
  reflecting transport canonical-JSON + actions-renamed
  changes.

Release plumbing
----------------

* pyproject.toml: version bumped 0.7.0 -> 0.7.6
* src/nullrun/__version__.py: __version__ = "0.7.6"
* CHANGELOG.md: full 0.7.6 entry covering additions,
  transport changes, metadata improvements

Tests pass locally (per session log) — pytest on Windows /
Python 3.14.2 is green.

* ci: fix PR #35 — fastapi dep + Transport._send_batch typo + coverage padding

PR #35 (release/0.7.6) failed all four CI jobs (test 3.10/3.11/3.12,
coverage, codecov/patch) on the same root cause + one latent bug
masked by it. This commit lands the fixes plus the last-mile tests
that bring coverage above the 82% threshold.

CI failure root
---------------

* tests/test_integrations_fastapi.py does from fastapi import ...
  at module top-level. CI installs only pip install -e '.[dev]',
  and fastapi was declared as an *optional* [fastapi] extra,
  NOT in [dev]. Pytest collection aborted with
  ModuleNotFoundError: No module named 'fastapi' → all 4 jobs red.
* Fix: add fastapi>=0.100,<1.0 to [dev]. Same precedent as
  langchain-core (already in [dev] for the same import-time
  contract: nullrun.instrumentation.langgraph is eager-imported
  from nullrun.decorators at collection time, so the test extras
  must cover the import chain).

Latent bug surfaced by the first fix
------------------------------------

The same PR refactored Transport._send_batch_with_retry_info to
route the /track/batch body through _signed_request_body for
canonical-JSON serialization (matching /gate and /execute). The two
sibling call sites use the module-level helper _signed_request_body
(no self.); this one used self._signed_request_body by typo.
Result: AttributeError on every batch flush, breaking 15 existing
tests across test_transport.py / test_track_batch_retry.py /
test_integration_contract.py / test_signal_safety.py. As long as
the fastapi collection error aborted pytest, this was hidden. Fixed
to _signed_request_body(...) with a docstring noting why it is
module-level and what the bug looked like.

Coverage padding (codecov/patch was failing on this too)
--------------------------------------------------------

Total coverage on the failing CI run was 81.98% — 0.02pp under the
fail-under=82 gate. After the two fixes above it would have
recovered to ~82.0% on the dot, so I added minimal tests for the
cheapest-to-cover gaps:

* tests/test_breaker_main.py (new) — covers the 5 statements in
  nullrun.breaker.__main__.main() (0% → 100%). The module
  exists so python -m nullrun.breaker exits cleanly instead of
  failing with No module named nullrun.breaker.__main__; the
  previous fix-mechanism was return 0 after a print, but no
  test was exercising it.
* tests/test_status.py — extends TestSummary with seven
  scenarios covering each conditional branch of NullRunStatus.summary()
  (organization_id, workflow_id, workflow_state != Normal,
  backend_reachable=False, ws_connected=False, recent_errors).
  status.py jumps 84.52% → 98.81%.
* tests/test_integrations_fastapi.py — four tests on
  _build_headers covering non-numeric, zero, negative, and
  resume_after (the WorkflowPausedException code path).
  integrations/fastapi.py jumps 90.22% → 94.57%.

After all three: TOTAL 81.98% → 82.46%, comfortably above the gate.

Verification
------------

* Local pytest: 997 passed, 13 skipped, 0 failed
  (Windows / Python 3.14.2, 8m47s — same env the original commit
  was validated in).
* python -m coverage report — 82.46%, no fail-under complaint.

* test: cover Phase 4.1 instrumentation — finish_reason + cache/reasoning/tools

Patch coverage on PR #35 was 62.38% against a 65% threshold (codecov
target 70% / threshold 5pp). The two biggest delta-holders against
master were auto.py (+286) and langgraph.py (+221), both dominated
by Phase 4.1 additions:

  * auto._normalize_finish_reason + _FINISH_REASON_MAP
  * auto._openai_extractor  second-tier fields (cache_read_tokens,
    cache_write_tokens, reasoning_tokens, finish_reason, tool_names)
  * auto._anthropic_extractor cache_read / cache_write
  * langgraph._safe_get_gen_message
  * langgraph._get_finish_reason (5-source fallback chain)
  * langgraph.extract_usage_from_response second-tier fields

These are pure / near-pure functions with no network or vendor SDK
calls. Coverage padding is cheap — pin the canonical wire shapes
once and the backend ingest contract gets a free live spec.

Local numbers:
  * auto.py        63.44% -> 64.01%   (file-level, +57 statements)
  * langgraph.py   78.50% -> 86.01%   (file-level, +32 statements)
  * TOTAL          82.46% -> 83.13%   (already above 82% gate)

41 tests, all green. Existing test_extractors.py and
test_langgraph_callback.py left untouched — these tests
deliberately target the Phase 4.1 fields (cache_read /
cache_write / reasoning / finish_reason / tool_names) that the
older tests didn't pin.

* fix(gate): forward real model + tools to /gate pre-flight (T4)

Pre-0.7.7 every SDK /gate call for any workflow with a budget was

hard-blocked because the runtime hard-coded the literal string

"budget-precheck" as the model. The backend's PolicyEvaluationGraph

treated any synthetic cost_limit rule with score > 0.8 as Block,

so the pricing lookup never landed on a real model and the rule

fired with the wrong score.

This commit:

* Adds nullrun.set_call_context(model=..., tools=[...]) plus

  get_call_model / get_call_tools helpers (and the underlying

  _call_model_var / _call_tools_var contextvars in

  nullrun.context).

* Wires the call context into check_workflow_budget: the /gate

  payload now carries the real model name (or None when unset)

  and the user-supplied tool list. tools=[] vs missing-None are

  distinguished on the wire per gate/internal.rs::check_tool_block.

* Transport.check forwards the tools key when set (it was

  silently dropped pre-fix).

* tests/conftest.py reset_runtime clears the new contextvars so

  a test's set_call_context(...) doesn't leak into the next

  test's wire payload.

* New tests/test_gate_real_path.py pins down the regression:

  default request allows a clean workflow, real block still

  honored, no policy-N residue on the wire, set_call_context

  flows into the body, no-context means no tools key, and the

  helpers are reachable from nullrun.*.

Bumps version to 0.7.7. No breaking changes - new helpers

default to None / empty so existing call sites keep working.

* release: 0.7.8 — fail-loud on deprecated surface

Two silent fail-OPEN footguns are converted to explicit
DeprecationWarning / RuntimeError so misconfigurations show up at
SDK init instead of being diagnosed from a missing proto trace.

Deprecated:

* NullRunRuntime.start_recording() and .stop_recording() now emit
  DeprecationWarning. They have been silent no-op stubs since
  Sprint 2.1 (0.4.0) — decision history is now on the backend
  dashboard at /control-center/decision-history. Both methods
  will be removed in 0.9.0.

* NULLRUN_USE_GRPC=1 now raises RuntimeError at SDK init instead
  of silently falling back to HTTP with an info log. gRPC is on
  the roadmap but not implemented; unset the env var to use HTTP.

Hardening (init path):

* Transport._post_auth_with_retry (new) — retry transient 503 / 504
  + network blips during /api/v1/auth/verify. Backend emits 503
  + Retry-After: 5 on transient DB errors (handlers.rs:11346-51).
  Pre-fix the first 503 surfaced as NR-A001 to the user as if the
  API key were bad. Three attempts, exponential backoff
  (0.5s → 1s → 2s), honors Retry-After when present. Auth-key
  failures (401) are NOT retried — a wrong key on attempt 1 is a
  wrong key on attempt 3.

Transport refactor:

* Transport._add_hmac_headers (new) — pulls the HMAC header
  construction out of _signed_request_body so /track/batch,
  /gate, /check, /execute all share one source of truth for
  Content-Type / X-Signature / X-Signature-Timestamp / X-API-Key
  / Authorization headers. HMAC formula unchanged.

* generate_hmac_signature + verify_hmac_signature accept str | bytes
  for body. Legacy str callers (and the FastAPI integration) keep
  working without an explicit .encode().

* actions_taken → actions on /track/batch response. Backend renamed
  BatchTrackResponse.actions_taken (debug names) → actions
  (ActionTaken structs with human-readable strings moved to
  messages). Read both keys for forward-compat.

Test updates:

* tests/test_framework_patches — alignment with retry + actions
  rename.
* tests/test_high_reliability_fixes — re-pinned for _post_auth_with_retry.
* tests/test_hmac_signing — expanded for str/bytes body + new
  _add_hmac_headers helper.
* tests/test_integration_contract — backend actions rename covered.
* tests/test_transport — retry semantics.

Bumps version to 0.7.8. No breaking changes for callers who don't
touch start_recording / stop_recording / NULLRUN_USE_GRPC.

* test(grpc): align test_grpc_removed with 0.7.8 NULLRUN_USE_GRPC contract

The 0.7.8 commit changed NULLRUN_USE_GRPC=1 from silent no-op +
INFO log to an explicit RuntimeError, but the regression test
in tests/test_grpc_removed.py still pinned the old behavior
(``test_nullrun_use_grpc_does_not_crash_init`` asserting
make_runtime() succeeded and an INFO line was logged).

CI on PR #38 failed on this test:

  FAILED tests/test_grpc_removed.py::TestGrpcRemoved
    ::test_nullrun_use_grpc_does_not_crash_init
  E   RuntimeError: NULLRUN_USE_GRPC is set but the gRPC
      transport is not yet implemented. ...

This commit updates the test to pin the new 0.7.8 contract:
the env var must raise RuntimeError, and the error message
must name the offending variable + point at the docs page.

The test is renamed from
``test_nullrun_use_grpc_does_not_crash_init`` to
``test_nullrun_use_grpc_raises_runtime_error`` so the test
name itself documents the new contract.

The module docstring (point 2 in the contract list) is
updated to say "raises RuntimeError" instead of "does NOT
crash init — it logs an INFO line and silently falls back
to HTTP". The 0.3.1 -> 0.7.8 evolution is documented in the
test docstring as a contract-evolution footnote for future
maintainers.

Imports: removed unused `import logging` and `caplog`
parameter (no longer asserting on log records); added
`import pytest` for `pytest.raises`.

No production-code change. No version bump. The fix is
self-contained to tests/test_grpc_removed.py.

* style(runtime): sort stdlib imports (ruff I001)

The 0.7.8 commit (fail-loud on deprecated surface) added
``import warnings`` mid-block in src/nullrun/runtime.py:34,
breaking alphabetical order:

    asyncio
    logging
    os
    warnings       <-- out of order
    threading
    time
    uuid

Ruff on PR #38 CI (Run ruff check src/) flagged it as I001.

Reorder to alphabetical:

    asyncio
    logging
    os
    threading
    time
    uuid
    warnings

Verified:
  * ruff check src/ -> All checks passed!
  * pytest tests/test_grpc_removed.py tests/test_runtime_branches.py
    -> 47 passed

No behavior change, no production logic touched. Pure lint fix.
maltsev-dev added a commit that referenced this pull request Aug 7, 2026
* release: 0.7.6 — FastAPI integration + user-facing message catalog

Additive patch on top of the 0.7.0 thin-client refactor. No
breaking changes.

Added
-----

* nullrun.integrations.fastapi — one-line FastAPI integration
  that turns every NullRunDecision / NullRunInfrastructureError
  thrown by @nullrun.protect endpoints into a clean JSON
  response with the right HTTP status code. No per-endpoint
  except blocks required.

  Response shape:
    {"error_code": "NR-B004",
     "user_message": "You've reached the usage limit...",
     "category": "decision"}

  HTTP status mapping:
    * NR-B004 (budget), NR-L001 (loop), NR-R001 (rate) -> 429
      with optional Retry-After
    * NR-T001 (tool blocked), NR-X001 (generic block) -> 403
    * NR-W003 (paused) -> 503 with Retry-After
    * NR-W002 (killed) -> 503; WorkflowKilledInterrupt is a
      BaseException subclass so Starlette's
      add_exception_handler refuses it — handled via ASGI
      middleware instead (hybrid pattern, documented in
      module docstring).
    * NullRunInfrastructureError subclasses -> 503 (our side,
      not user's).

* nullrun.messages — default user-facing message catalog.
  Every NR-* error code has an English default message owned
  by NULLRUN, not customer code. Customer Support Bots hitting
  a budget cap show the same wording across every NullRun-backed
  application.
    * format_user_message(exc) — render exception as user-facing
      string
    * set_user_message(code, text) — per-process override for
      branded variants
    * get_user_message(code) — raw lookup
    * reset_overrides() — clear all overrides (for tests)

Changed
-------

* Transport._send_batch canonical JSON serialization — route the
  /track/batch body through _signed_request_body for consistent
  compact-separator serialization. HMAC itself is unaffected,
  but consistent serialization removes a special-case from the
  wire-format contract tests.

* Transport._send_batch actions response handling — backend
  renamed BatchTrackResponse.actions_taken (debug names) ->
  BatchTrackResponse.actions (ActionTaken structs). Read both
  for forward-compat; per-element try/except so one malformed
  entry doesn't abort the whole loop.

* pyproject.toml metadata — long-form description with search
  keywords, Maintainer: populated via maintainers=[...],
  expanded classifiers (Linux / Windows / macOS, Python 3.13,
  CPython, Security / AI / WWW/HTTP topics), project URL
  expander.

Tests
-----

* tests/test_messages.py (new, 282 lines) — catalog
  completeness (every NR-* code has a default message),
  override / reset behavior, render path.
* tests/test_integrations_fastapi.py (new, 289 lines) — HTTP
  status mapping per error code, response shape, ASGI
  middleware path for WorkflowKilledInterrupt, hybrid
  composition.
* tests/test_decision_split.py (new, 199 lines) — pins the
  decision / infrastructure error split.
* Updates to tests/test_runtime.py, tests/test_extractors.py
  reflecting transport canonical-JSON + actions-renamed
  changes.

Release plumbing
----------------

* pyproject.toml: version bumped 0.7.0 -> 0.7.6
* src/nullrun/__version__.py: __version__ = "0.7.6"
* CHANGELOG.md: full 0.7.6 entry covering additions,
  transport changes, metadata improvements

Tests pass locally (per session log) — pytest on Windows /
Python 3.14.2 is green.

* ci: fix PR #35 — fastapi dep + Transport._send_batch typo + coverage padding

PR #35 (release/0.7.6) failed all four CI jobs (test 3.10/3.11/3.12,
coverage, codecov/patch) on the same root cause + one latent bug
masked by it. This commit lands the fixes plus the last-mile tests
that bring coverage above the 82% threshold.

CI failure root
---------------

* tests/test_integrations_fastapi.py does from fastapi import ...
  at module top-level. CI installs only pip install -e '.[dev]',
  and fastapi was declared as an *optional* [fastapi] extra,
  NOT in [dev]. Pytest collection aborted with
  ModuleNotFoundError: No module named 'fastapi' → all 4 jobs red.
* Fix: add fastapi>=0.100,<1.0 to [dev]. Same precedent as
  langchain-core (already in [dev] for the same import-time
  contract: nullrun.instrumentation.langgraph is eager-imported
  from nullrun.decorators at collection time, so the test extras
  must cover the import chain).

Latent bug surfaced by the first fix
------------------------------------

The same PR refactored Transport._send_batch_with_retry_info to
route the /track/batch body through _signed_request_body for
canonical-JSON serialization (matching /gate and /execute). The two
sibling call sites use the module-level helper _signed_request_body
(no self.); this one used self._signed_request_body by typo.
Result: AttributeError on every batch flush, breaking 15 existing
tests across test_transport.py / test_track_batch_retry.py /
test_integration_contract.py / test_signal_safety.py. As long as
the fastapi collection error aborted pytest, this was hidden. Fixed
to _signed_request_body(...) with a docstring noting why it is
module-level and what the bug looked like.

Coverage padding (codecov/patch was failing on this too)
--------------------------------------------------------

Total coverage on the failing CI run was 81.98% — 0.02pp under the
fail-under=82 gate. After the two fixes above it would have
recovered to ~82.0% on the dot, so I added minimal tests for the
cheapest-to-cover gaps:

* tests/test_breaker_main.py (new) — covers the 5 statements in
  nullrun.breaker.__main__.main() (0% → 100%). The module
  exists so python -m nullrun.breaker exits cleanly instead of
  failing with No module named nullrun.breaker.__main__; the
  previous fix-mechanism was return 0 after a print, but no
  test was exercising it.
* tests/test_status.py — extends TestSummary with seven
  scenarios covering each conditional branch of NullRunStatus.summary()
  (organization_id, workflow_id, workflow_state != Normal,
  backend_reachable=False, ws_connected=False, recent_errors).
  status.py jumps 84.52% → 98.81%.
* tests/test_integrations_fastapi.py — four tests on
  _build_headers covering non-numeric, zero, negative, and
  resume_after (the WorkflowPausedException code path).
  integrations/fastapi.py jumps 90.22% → 94.57%.

After all three: TOTAL 81.98% → 82.46%, comfortably above the gate.

Verification
------------

* Local pytest: 997 passed, 13 skipped, 0 failed
  (Windows / Python 3.14.2, 8m47s — same env the original commit
  was validated in).
* python -m coverage report — 82.46%, no fail-under complaint.

* test: cover Phase 4.1 instrumentation — finish_reason + cache/reasoning/tools

Patch coverage on PR #35 was 62.38% against a 65% threshold (codecov
target 70% / threshold 5pp). The two biggest delta-holders against
master were auto.py (+286) and langgraph.py (+221), both dominated
by Phase 4.1 additions:

  * auto._normalize_finish_reason + _FINISH_REASON_MAP
  * auto._openai_extractor  second-tier fields (cache_read_tokens,
    cache_write_tokens, reasoning_tokens, finish_reason, tool_names)
  * auto._anthropic_extractor cache_read / cache_write
  * langgraph._safe_get_gen_message
  * langgraph._get_finish_reason (5-source fallback chain)
  * langgraph.extract_usage_from_response second-tier fields

These are pure / near-pure functions with no network or vendor SDK
calls. Coverage padding is cheap — pin the canonical wire shapes
once and the backend ingest contract gets a free live spec.

Local numbers:
  * auto.py        63.44% -> 64.01%   (file-level, +57 statements)
  * langgraph.py   78.50% -> 86.01%   (file-level, +32 statements)
  * TOTAL          82.46% -> 83.13%   (already above 82% gate)

41 tests, all green. Existing test_extractors.py and
test_langgraph_callback.py left untouched — these tests
deliberately target the Phase 4.1 fields (cache_read /
cache_write / reasoning / finish_reason / tool_names) that the
older tests didn't pin.

* fix(gate): forward real model + tools to /gate pre-flight (T4)

Pre-0.7.7 every SDK /gate call for any workflow with a budget was

hard-blocked because the runtime hard-coded the literal string

"budget-precheck" as the model. The backend's PolicyEvaluationGraph

treated any synthetic cost_limit rule with score > 0.8 as Block,

so the pricing lookup never landed on a real model and the rule

fired with the wrong score.

This commit:

* Adds nullrun.set_call_context(model=..., tools=[...]) plus

  get_call_model / get_call_tools helpers (and the underlying

  _call_model_var / _call_tools_var contextvars in

  nullrun.context).

* Wires the call context into check_workflow_budget: the /gate

  payload now carries the real model name (or None when unset)

  and the user-supplied tool list. tools=[] vs missing-None are

  distinguished on the wire per gate/internal.rs::check_tool_block.

* Transport.check forwards the tools key when set (it was

  silently dropped pre-fix).

* tests/conftest.py reset_runtime clears the new contextvars so

  a test's set_call_context(...) doesn't leak into the next

  test's wire payload.

* New tests/test_gate_real_path.py pins down the regression:

  default request allows a clean workflow, real block still

  honored, no policy-N residue on the wire, set_call_context

  flows into the body, no-context means no tools key, and the

  helpers are reachable from nullrun.*.

Bumps version to 0.7.7. No breaking changes - new helpers

default to None / empty so existing call sites keep working.

* release: 0.7.8 — fail-loud on deprecated surface

Two silent fail-OPEN footguns are converted to explicit
DeprecationWarning / RuntimeError so misconfigurations show up at
SDK init instead of being diagnosed from a missing proto trace.

Deprecated:

* NullRunRuntime.start_recording() and .stop_recording() now emit
  DeprecationWarning. They have been silent no-op stubs since
  Sprint 2.1 (0.4.0) — decision history is now on the backend
  dashboard at /control-center/decision-history. Both methods
  will be removed in 0.9.0.

* NULLRUN_USE_GRPC=1 now raises RuntimeError at SDK init instead
  of silently falling back to HTTP with an info log. gRPC is on
  the roadmap but not implemented; unset the env var to use HTTP.

Hardening (init path):

* Transport._post_auth_with_retry (new) — retry transient 503 / 504
  + network blips during /api/v1/auth/verify. Backend emits 503
  + Retry-After: 5 on transient DB errors (handlers.rs:11346-51).
  Pre-fix the first 503 surfaced as NR-A001 to the user as if the
  API key were bad. Three attempts, exponential backoff
  (0.5s → 1s → 2s), honors Retry-After when present. Auth-key
  failures (401) are NOT retried — a wrong key on attempt 1 is a
  wrong key on attempt 3.

Transport refactor:

* Transport._add_hmac_headers (new) — pulls the HMAC header
  construction out of _signed_request_body so /track/batch,
  /gate, /check, /execute all share one source of truth for
  Content-Type / X-Signature / X-Signature-Timestamp / X-API-Key
  / Authorization headers. HMAC formula unchanged.

* generate_hmac_signature + verify_hmac_signature accept str | bytes
  for body. Legacy str callers (and the FastAPI integration) keep
  working without an explicit .encode().

* actions_taken → actions on /track/batch response. Backend renamed
  BatchTrackResponse.actions_taken (debug names) → actions
  (ActionTaken structs with human-readable strings moved to
  messages). Read both keys for forward-compat.

Test updates:

* tests/test_framework_patches — alignment with retry + actions
  rename.
* tests/test_high_reliability_fixes — re-pinned for _post_auth_with_retry.
* tests/test_hmac_signing — expanded for str/bytes body + new
  _add_hmac_headers helper.
* tests/test_integration_contract — backend actions rename covered.
* tests/test_transport — retry semantics.

Bumps version to 0.7.8. No breaking changes for callers who don't
touch start_recording / stop_recording / NULLRUN_USE_GRPC.

* test(grpc): align test_grpc_removed with 0.7.8 NULLRUN_USE_GRPC contract

The 0.7.8 commit changed NULLRUN_USE_GRPC=1 from silent no-op +
INFO log to an explicit RuntimeError, but the regression test
in tests/test_grpc_removed.py still pinned the old behavior
(``test_nullrun_use_grpc_does_not_crash_init`` asserting
make_runtime() succeeded and an INFO line was logged).

CI on PR #38 failed on this test:

  FAILED tests/test_grpc_removed.py::TestGrpcRemoved
    ::test_nullrun_use_grpc_does_not_crash_init
  E   RuntimeError: NULLRUN_USE_GRPC is set but the gRPC
      transport is not yet implemented. ...

This commit updates the test to pin the new 0.7.8 contract:
the env var must raise RuntimeError, and the error message
must name the offending variable + point at the docs page.

The test is renamed from
``test_nullrun_use_grpc_does_not_crash_init`` to
``test_nullrun_use_grpc_raises_runtime_error`` so the test
name itself documents the new contract.

The module docstring (point 2 in the contract list) is
updated to say "raises RuntimeError" instead of "does NOT
crash init — it logs an INFO line and silently falls back
to HTTP". The 0.3.1 -> 0.7.8 evolution is documented in the
test docstring as a contract-evolution footnote for future
maintainers.

Imports: removed unused `import logging` and `caplog`
parameter (no longer asserting on log records); added
`import pytest` for `pytest.raises`.

No production-code change. No version bump. The fix is
self-contained to tests/test_grpc_removed.py.

* style(runtime): sort stdlib imports (ruff I001)

The 0.7.8 commit (fail-loud on deprecated surface) added
``import warnings`` mid-block in src/nullrun/runtime.py:34,
breaking alphabetical order:

    asyncio
    logging
    os
    warnings       <-- out of order
    threading
    time
    uuid

Ruff on PR #38 CI (Run ruff check src/) flagged it as I001.

Reorder to alphabetical:

    asyncio
    logging
    os
    threading
    time
    uuid
    warnings

Verified:
  * ruff check src/ -> All checks passed!
  * pytest tests/test_grpc_removed.py tests/test_runtime_branches.py
    -> 47 passed

No behavior change, no production logic touched. Pure lint fix.

* release: 0.8.0 — SDK wire-format audit (model/provider extraction)

Closes a class of silent-fail-OPEN path that was sending
model=None or model="unknown" on /track for many LLM-vendor
paths. Every such event cost the backend a model_pricing
lookup that returned no row, fell through to DEFAULT_RATE
(~$30/M), and emitted a fallback warning the operator
couldn't reproduce because the offending observation was
buried in another package's telemetry.

No public-API break. No behavior change for callers whose
instrumentation already populates model correctly. Pure
wire-payload hygiene.

runtime.py — track():

* Strips None values from the wire payload (pre-0.8.0
  forwarded every key except _WIRE_STRIP_FIELDS, including
  keys whose value was None). Putting {"model": null} on
  the wire triggered backend unwrap_or("default") and a
  fallback warning. Dropping None keeps the diagnostic
  signal loud (the new WARN below fires on missing-key,
  which is what we want operators to see) instead of
  silent (the JSON-null case).

* Adds logger.warning("track(): llm_call event missing
  'model' field — backend will fall back to DEFAULT_RATE.
  event=...") — the single signal an operator needs to
  reproduce "which observation produced an llm_call
  without model set". Activated only for llm_call; other
  event types are silent.

instrumentation/langgraph.py — NullRunCallback.on_llm_end:

* New _extract_model_from_response + _extract_provider_from_response
  helpers (mirror _get_finish_reason's best-effort
  pattern). Fallback chain: invocation_params → response
  metadata → AIMessage response_metadata → llm_output →
  direct attribute. "unknown" is now a true last resort,
  not the common case.

instrumentation/llama_index.py:

* extract_from_event fallback chain: event.response.model
  → event.response.raw.model → usage['model']. Mock
  providers and adapter-style ChatResponse now ship a
  real model id.

instrumentation/autogen.py:

* on_messages fallback chain: self.model → result.model.
  OpenAI's response carries the actual model id (may
  differ from request if the server resolved an alias).

instrumentation/auto.py — _emit_from_span (openai-agents):

* span model fallback chain: span['model'] →
  usage['model'] → span['response_metadata']['model_name'].
  Some custom tracer configs leave span['model'] empty;
  the other two sources usually have it.

  Sets model on the event only when we have a real value
  (empty/None is dropped — relies on the new None-strip
  in track() to keep the operator warning loud).

Bumps version to 0.8.0. No breaking changes for callers
who don't touch the wire path directly.
maltsev-dev added a commit that referenced this pull request Aug 7, 2026
…40)

* release: 0.7.6 — FastAPI integration + user-facing message catalog

Additive patch on top of the 0.7.0 thin-client refactor. No
breaking changes.

Added
-----

* nullrun.integrations.fastapi — one-line FastAPI integration
  that turns every NullRunDecision / NullRunInfrastructureError
  thrown by @nullrun.protect endpoints into a clean JSON
  response with the right HTTP status code. No per-endpoint
  except blocks required.

  Response shape:
    {"error_code": "NR-B004",
     "user_message": "You've reached the usage limit...",
     "category": "decision"}

  HTTP status mapping:
    * NR-B004 (budget), NR-L001 (loop), NR-R001 (rate) -> 429
      with optional Retry-After
    * NR-T001 (tool blocked), NR-X001 (generic block) -> 403
    * NR-W003 (paused) -> 503 with Retry-After
    * NR-W002 (killed) -> 503; WorkflowKilledInterrupt is a
      BaseException subclass so Starlette's
      add_exception_handler refuses it — handled via ASGI
      middleware instead (hybrid pattern, documented in
      module docstring).
    * NullRunInfrastructureError subclasses -> 503 (our side,
      not user's).

* nullrun.messages — default user-facing message catalog.
  Every NR-* error code has an English default message owned
  by NULLRUN, not customer code. Customer Support Bots hitting
  a budget cap show the same wording across every NullRun-backed
  application.
    * format_user_message(exc) — render exception as user-facing
      string
    * set_user_message(code, text) — per-process override for
      branded variants
    * get_user_message(code) — raw lookup
    * reset_overrides() — clear all overrides (for tests)

Changed
-------

* Transport._send_batch canonical JSON serialization — route the
  /track/batch body through _signed_request_body for consistent
  compact-separator serialization. HMAC itself is unaffected,
  but consistent serialization removes a special-case from the
  wire-format contract tests.

* Transport._send_batch actions response handling — backend
  renamed BatchTrackResponse.actions_taken (debug names) ->
  BatchTrackResponse.actions (ActionTaken structs). Read both
  for forward-compat; per-element try/except so one malformed
  entry doesn't abort the whole loop.

* pyproject.toml metadata — long-form description with search
  keywords, Maintainer: populated via maintainers=[...],
  expanded classifiers (Linux / Windows / macOS, Python 3.13,
  CPython, Security / AI / WWW/HTTP topics), project URL
  expander.

Tests
-----

* tests/test_messages.py (new, 282 lines) — catalog
  completeness (every NR-* code has a default message),
  override / reset behavior, render path.
* tests/test_integrations_fastapi.py (new, 289 lines) — HTTP
  status mapping per error code, response shape, ASGI
  middleware path for WorkflowKilledInterrupt, hybrid
  composition.
* tests/test_decision_split.py (new, 199 lines) — pins the
  decision / infrastructure error split.
* Updates to tests/test_runtime.py, tests/test_extractors.py
  reflecting transport canonical-JSON + actions-renamed
  changes.

Release plumbing
----------------

* pyproject.toml: version bumped 0.7.0 -> 0.7.6
* src/nullrun/__version__.py: __version__ = "0.7.6"
* CHANGELOG.md: full 0.7.6 entry covering additions,
  transport changes, metadata improvements

Tests pass locally (per session log) — pytest on Windows /
Python 3.14.2 is green.

* ci: fix PR #35 — fastapi dep + Transport._send_batch typo + coverage padding

PR #35 (release/0.7.6) failed all four CI jobs (test 3.10/3.11/3.12,
coverage, codecov/patch) on the same root cause + one latent bug
masked by it. This commit lands the fixes plus the last-mile tests
that bring coverage above the 82% threshold.

CI failure root
---------------

* tests/test_integrations_fastapi.py does from fastapi import ...
  at module top-level. CI installs only pip install -e '.[dev]',
  and fastapi was declared as an *optional* [fastapi] extra,
  NOT in [dev]. Pytest collection aborted with
  ModuleNotFoundError: No module named 'fastapi' → all 4 jobs red.
* Fix: add fastapi>=0.100,<1.0 to [dev]. Same precedent as
  langchain-core (already in [dev] for the same import-time
  contract: nullrun.instrumentation.langgraph is eager-imported
  from nullrun.decorators at collection time, so the test extras
  must cover the import chain).

Latent bug surfaced by the first fix
------------------------------------

The same PR refactored Transport._send_batch_with_retry_info to
route the /track/batch body through _signed_request_body for
canonical-JSON serialization (matching /gate and /execute). The two
sibling call sites use the module-level helper _signed_request_body
(no self.); this one used self._signed_request_body by typo.
Result: AttributeError on every batch flush, breaking 15 existing
tests across test_transport.py / test_track_batch_retry.py /
test_integration_contract.py / test_signal_safety.py. As long as
the fastapi collection error aborted pytest, this was hidden. Fixed
to _signed_request_body(...) with a docstring noting why it is
module-level and what the bug looked like.

Coverage padding (codecov/patch was failing on this too)
--------------------------------------------------------

Total coverage on the failing CI run was 81.98% — 0.02pp under the
fail-under=82 gate. After the two fixes above it would have
recovered to ~82.0% on the dot, so I added minimal tests for the
cheapest-to-cover gaps:

* tests/test_breaker_main.py (new) — covers the 5 statements in
  nullrun.breaker.__main__.main() (0% → 100%). The module
  exists so python -m nullrun.breaker exits cleanly instead of
  failing with No module named nullrun.breaker.__main__; the
  previous fix-mechanism was return 0 after a print, but no
  test was exercising it.
* tests/test_status.py — extends TestSummary with seven
  scenarios covering each conditional branch of NullRunStatus.summary()
  (organization_id, workflow_id, workflow_state != Normal,
  backend_reachable=False, ws_connected=False, recent_errors).
  status.py jumps 84.52% → 98.81%.
* tests/test_integrations_fastapi.py — four tests on
  _build_headers covering non-numeric, zero, negative, and
  resume_after (the WorkflowPausedException code path).
  integrations/fastapi.py jumps 90.22% → 94.57%.

After all three: TOTAL 81.98% → 82.46%, comfortably above the gate.

Verification
------------

* Local pytest: 997 passed, 13 skipped, 0 failed
  (Windows / Python 3.14.2, 8m47s — same env the original commit
  was validated in).
* python -m coverage report — 82.46%, no fail-under complaint.

* test: cover Phase 4.1 instrumentation — finish_reason + cache/reasoning/tools

Patch coverage on PR #35 was 62.38% against a 65% threshold (codecov
target 70% / threshold 5pp). The two biggest delta-holders against
master were auto.py (+286) and langgraph.py (+221), both dominated
by Phase 4.1 additions:

  * auto._normalize_finish_reason + _FINISH_REASON_MAP
  * auto._openai_extractor  second-tier fields (cache_read_tokens,
    cache_write_tokens, reasoning_tokens, finish_reason, tool_names)
  * auto._anthropic_extractor cache_read / cache_write
  * langgraph._safe_get_gen_message
  * langgraph._get_finish_reason (5-source fallback chain)
  * langgraph.extract_usage_from_response second-tier fields

These are pure / near-pure functions with no network or vendor SDK
calls. Coverage padding is cheap — pin the canonical wire shapes
once and the backend ingest contract gets a free live spec.

Local numbers:
  * auto.py        63.44% -> 64.01%   (file-level, +57 statements)
  * langgraph.py   78.50% -> 86.01%   (file-level, +32 statements)
  * TOTAL          82.46% -> 83.13%   (already above 82% gate)

41 tests, all green. Existing test_extractors.py and
test_langgraph_callback.py left untouched — these tests
deliberately target the Phase 4.1 fields (cache_read /
cache_write / reasoning / finish_reason / tool_names) that the
older tests didn't pin.

* fix(gate): forward real model + tools to /gate pre-flight (T4)

Pre-0.7.7 every SDK /gate call for any workflow with a budget was

hard-blocked because the runtime hard-coded the literal string

"budget-precheck" as the model. The backend's PolicyEvaluationGraph

treated any synthetic cost_limit rule with score > 0.8 as Block,

so the pricing lookup never landed on a real model and the rule

fired with the wrong score.

This commit:

* Adds nullrun.set_call_context(model=..., tools=[...]) plus

  get_call_model / get_call_tools helpers (and the underlying

  _call_model_var / _call_tools_var contextvars in

  nullrun.context).

* Wires the call context into check_workflow_budget: the /gate

  payload now carries the real model name (or None when unset)

  and the user-supplied tool list. tools=[] vs missing-None are

  distinguished on the wire per gate/internal.rs::check_tool_block.

* Transport.check forwards the tools key when set (it was

  silently dropped pre-fix).

* tests/conftest.py reset_runtime clears the new contextvars so

  a test's set_call_context(...) doesn't leak into the next

  test's wire payload.

* New tests/test_gate_real_path.py pins down the regression:

  default request allows a clean workflow, real block still

  honored, no policy-N residue on the wire, set_call_context

  flows into the body, no-context means no tools key, and the

  helpers are reachable from nullrun.*.

Bumps version to 0.7.7. No breaking changes - new helpers

default to None / empty so existing call sites keep working.

* release: 0.7.8 — fail-loud on deprecated surface

Two silent fail-OPEN footguns are converted to explicit
DeprecationWarning / RuntimeError so misconfigurations show up at
SDK init instead of being diagnosed from a missing proto trace.

Deprecated:

* NullRunRuntime.start_recording() and .stop_recording() now emit
  DeprecationWarning. They have been silent no-op stubs since
  Sprint 2.1 (0.4.0) — decision history is now on the backend
  dashboard at /control-center/decision-history. Both methods
  will be removed in 0.9.0.

* NULLRUN_USE_GRPC=1 now raises RuntimeError at SDK init instead
  of silently falling back to HTTP with an info log. gRPC is on
  the roadmap but not implemented; unset the env var to use HTTP.

Hardening (init path):

* Transport._post_auth_with_retry (new) — retry transient 503 / 504
  + network blips during /api/v1/auth/verify. Backend emits 503
  + Retry-After: 5 on transient DB errors (handlers.rs:11346-51).
  Pre-fix the first 503 surfaced as NR-A001 to the user as if the
  API key were bad. Three attempts, exponential backoff
  (0.5s → 1s → 2s), honors Retry-After when present. Auth-key
  failures (401) are NOT retried — a wrong key on attempt 1 is a
  wrong key on attempt 3.

Transport refactor:

* Transport._add_hmac_headers (new) — pulls the HMAC header
  construction out of _signed_request_body so /track/batch,
  /gate, /check, /execute all share one source of truth for
  Content-Type / X-Signature / X-Signature-Timestamp / X-API-Key
  / Authorization headers. HMAC formula unchanged.

* generate_hmac_signature + verify_hmac_signature accept str | bytes
  for body. Legacy str callers (and the FastAPI integration) keep
  working without an explicit .encode().

* actions_taken → actions on /track/batch response. Backend renamed
  BatchTrackResponse.actions_taken (debug names) → actions
  (ActionTaken structs with human-readable strings moved to
  messages). Read both keys for forward-compat.

Test updates:

* tests/test_framework_patches — alignment with retry + actions
  rename.
* tests/test_high_reliability_fixes — re-pinned for _post_auth_with_retry.
* tests/test_hmac_signing — expanded for str/bytes body + new
  _add_hmac_headers helper.
* tests/test_integration_contract — backend actions rename covered.
* tests/test_transport — retry semantics.

Bumps version to 0.7.8. No breaking changes for callers who don't
touch start_recording / stop_recording / NULLRUN_USE_GRPC.

* test(grpc): align test_grpc_removed with 0.7.8 NULLRUN_USE_GRPC contract

The 0.7.8 commit changed NULLRUN_USE_GRPC=1 from silent no-op +
INFO log to an explicit RuntimeError, but the regression test
in tests/test_grpc_removed.py still pinned the old behavior
(``test_nullrun_use_grpc_does_not_crash_init`` asserting
make_runtime() succeeded and an INFO line was logged).

CI on PR #38 failed on this test:

  FAILED tests/test_grpc_removed.py::TestGrpcRemoved
    ::test_nullrun_use_grpc_does_not_crash_init
  E   RuntimeError: NULLRUN_USE_GRPC is set but the gRPC
      transport is not yet implemented. ...

This commit updates the test to pin the new 0.7.8 contract:
the env var must raise RuntimeError, and the error message
must name the offending variable + point at the docs page.

The test is renamed from
``test_nullrun_use_grpc_does_not_crash_init`` to
``test_nullrun_use_grpc_raises_runtime_error`` so the test
name itself documents the new contract.

The module docstring (point 2 in the contract list) is
updated to say "raises RuntimeError" instead of "does NOT
crash init — it logs an INFO line and silently falls back
to HTTP". The 0.3.1 -> 0.7.8 evolution is documented in the
test docstring as a contract-evolution footnote for future
maintainers.

Imports: removed unused `import logging` and `caplog`
parameter (no longer asserting on log records); added
`import pytest` for `pytest.raises`.

No production-code change. No version bump. The fix is
self-contained to tests/test_grpc_removed.py.

* style(runtime): sort stdlib imports (ruff I001)

The 0.7.8 commit (fail-loud on deprecated surface) added
``import warnings`` mid-block in src/nullrun/runtime.py:34,
breaking alphabetical order:

    asyncio
    logging
    os
    warnings       <-- out of order
    threading
    time
    uuid

Ruff on PR #38 CI (Run ruff check src/) flagged it as I001.

Reorder to alphabetical:

    asyncio
    logging
    os
    threading
    time
    uuid
    warnings

Verified:
  * ruff check src/ -> All checks passed!
  * pytest tests/test_grpc_removed.py tests/test_runtime_branches.py
    -> 47 passed

No behavior change, no production logic touched. Pure lint fix.

* release: 0.8.0 — SDK wire-format audit (model/provider extraction)

Closes a class of silent-fail-OPEN path that was sending
model=None or model="unknown" on /track for many LLM-vendor
paths. Every such event cost the backend a model_pricing
lookup that returned no row, fell through to DEFAULT_RATE
(~$30/M), and emitted a fallback warning the operator
couldn't reproduce because the offending observation was
buried in another package's telemetry.

No public-API break. No behavior change for callers whose
instrumentation already populates model correctly. Pure
wire-payload hygiene.

runtime.py — track():

* Strips None values from the wire payload (pre-0.8.0
  forwarded every key except _WIRE_STRIP_FIELDS, including
  keys whose value was None). Putting {"model": null} on
  the wire triggered backend unwrap_or("default") and a
  fallback warning. Dropping None keeps the diagnostic
  signal loud (the new WARN below fires on missing-key,
  which is what we want operators to see) instead of
  silent (the JSON-null case).

* Adds logger.warning("track(): llm_call event missing
  'model' field — backend will fall back to DEFAULT_RATE.
  event=...") — the single signal an operator needs to
  reproduce "which observation produced an llm_call
  without model set". Activated only for llm_call; other
  event types are silent.

instrumentation/langgraph.py — NullRunCallback.on_llm_end:

* New _extract_model_from_response + _extract_provider_from_response
  helpers (mirror _get_finish_reason's best-effort
  pattern). Fallback chain: invocation_params → response
  metadata → AIMessage response_metadata → llm_output →
  direct attribute. "unknown" is now a true last resort,
  not the common case.

instrumentation/llama_index.py:

* extract_from_event fallback chain: event.response.model
  → event.response.raw.model → usage['model']. Mock
  providers and adapter-style ChatResponse now ship a
  real model id.

instrumentation/autogen.py:

* on_messages fallback chain: self.model → result.model.
  OpenAI's response carries the actual model id (may
  differ from request if the server resolved an alias).

instrumentation/auto.py — _emit_from_span (openai-agents):

* span model fallback chain: span['model'] →
  usage['model'] → span['response_metadata']['model_name'].
  Some custom tracer configs leave span['model'] empty;
  the other two sources usually have it.

  Sets model on the event only when we have a real value
  (empty/None is dropped — relies on the new None-strip
  in track() to keep the operator warning loud).

Bumps version to 0.8.0. No breaking changes for callers
who don't touch the wire path directly.

* fix: 0.8.2 — coverage wire-shape (metadata nesting) + model fallback

Two coordinated fixes from the 0.8.0 wire-format audit:

1. Coverage counters under metadata
   - src/nullrun/runtime.py: track_coverage() emits seen/tracked/
     streaming_skipped dicts under event.metadata instead of the
     top level. SdkTrackRequest uses explicit fields with no
     #[serde(flatten)] catchall, so top-level keys were silently
     dropped by serde and the dashboard's last_coverage_pct was
     permanently null.
   - tests/test_coverage_report.py: pin the wire shape (regression
     test).

2. Model name extraction fallback (Issue 2)
   - src/nullrun/instrumentation/auto.py: when the response body
     extractor returns None for model (OpenAI Responses API,
     Anthropic streaming edge cases), fall back to the model
     string the user embedded in the request body via
     ChatOpenAI(model='gpt-4.1-mini'). Without this, every such
     call was zero-billed (backend unwrap_or('default') +
     DEFAULT_RATE ≈ $0/call).
   - tests/test_model_fallback.py: unit-test the helper.

3. Backend batch response schema contract tests
   - tests/test_batch_response_parsing.py: pin the post-2026-06-27
     BatchTrackResponse shape (actions: Vec<ActionTaken>,
     messages: Vec<String>) and document that the legacy
     actions_taken: Vec<String> field is intentionally dropped in
     0.8.0.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant