release/v0.3.1: Approval-rate fix + accurate limitations - #3
Merged
Merged
Conversation
Fixes:
- get_tool_metrics SQL was querying `$.decision` on tool_result events,
which Claude Code never emits (real attribute is `decision_type`). All
approval rates silently fell back to 100%. Rewrote the query to:
- source accepts from tool_result.decision_type (values: `accept`)
- source rejects from tool_decision.decision (the only place rejects
surface, since rejected calls never produce a tool_result)
- reconcile MCP rejections via tool_use_id join, so they attribute to
the proper mcp__server__tool name instead of the generic mcp_tool
that tool_decision events carry.
- Added missing Claude Code built-ins (TaskCreate, TaskUpdate, TaskOutput,
TaskStop, TaskList, TaskGet, ToolSearch, TestRead) — they were rendering
as `mcp` in the TYPE column.
Tests (the missing verification layer):
- test_approval_rate_from_real_event_shape — round-trips realistic Claude
Code event shapes through duckdb-rs; asserts the tool_use_id join works
and no phantom mcp_tool row leaks.
- test_app_approval_rate_end_to_end — full App → storage → tool_metrics
pipeline; ExitPlanMode 25% (1/4), Bash 67% (2/3).
Docs:
- README limitations: all 4 rewritten against reality. MCP names were
fixed upstream 2026-03-25 (anthropic/claude-code#17046). Context window
is now scraped locally. Codex telemetry was fixed 2026-02-28
(openai/codex#12913). Approval rate now explained correctly.
- README providers table: OpenAI Codex CLI bumped Partial → Full.
Verification: 272 tests pass, clippy clean under -D warnings, fmt clean.
Real DuckDB data confirms new query: ExitPlanMode 3%, Bash 99%,
AskUserQuestion 83% (matches actual usage patterns).
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #3 +/- ##
===========================================
+ Coverage 59.62% 73.92% +14.30%
===========================================
Files 25 25
Lines 4604 4698 +94
===========================================
+ Hits 2745 3473 +728
+ Misses 1859 1225 -634 ☔ View full report in Codecov by Sentry. 🚀 New features to boost your workflow:
|
Mirrors the install_statusline_hook_in pattern already used in src/config/mod.rs. Each Provider trait impl's ensure_configured() now delegates to a public inherent method that takes the settings.json path explicitly, so tests can drive each provider against a tempdir instead of touching the user's real home directory. No behavior change. Targets the providers/* coverage gap (currently 18-56%) by making the auto-config branches testable.
Provider tests (21 new tests, ~+330 covered regions in providers/*): - tests/providers_claude_code_test.rs (6 tests) - tests/providers_gemini_test.rs (5 tests) - tests/providers_qwen_test.rs (5 tests) - tests/providers_copilot_test.rs (5 tests) Each exercises: fresh-dir create, idempotent re-run, preserves unrelated keys, .bak created only on modification, env/telemetry block merging, parse-error propagation. Storage SQL edge cases (8 new tests, ~+125 covered regions): - Time-filter coverage for every getter that takes Option<DateTime<Utc>> (get_tool_metrics, get_api_metrics, get_distinct_sessions, get_distinct_service_names, get_compaction_stats). - Empty-DB safety test exercising every getter (catches the silent-panic class). - Token-rate sparkline boundary-row regression (the off-by-one fix). - Multi-bucket distribution sanity check. Bug fixed along the way: - get_distinct_sessions used '$.session.id' (unquoted) which DuckDB interpreted as path traversal session->id, always returning NULL against real telemetry where the key is literally 'session.id'. Fixed to '$."session.id"' matching the pattern already used for service.name. The test that surfaced this would have caught the bug at original review time. Total: 299 tests passing (272 -> 299, +27).
Drives the axum router via tower::ServiceExt::oneshot — no real TCP socket, no port races. JSON OTLP payloads (the parser falls back from protobuf to JSON, so JSON exercises the same downstream path). Coverage of the receiver wiring: - logs endpoint: tool_result and tool_decision events both persist and produce the correct ToolMetrics shape (regression for the approval rate bug). - metrics endpoint: token usage and cost both reach the right storage recorder. - traces endpoint: accepts but no-ops (current behavior documented). - compaction events flow into get_compaction_stats. - malformed JSON returns 200 with no storage state change (current swallow-and-log behavior documented). Adds tower 0.5 as a dev-dependency (already in the tree via axum, but needs explicit declaration for ServiceExt to be in scope). Refactors src/otlp/mod.rs to expose pub fn build_router(storage) so tests can share the same Router construction as start_receiver.
Caught real bug: quota panel only allocated 3 vertical rows, so the 7d quota row was clipped (only the 5h gauge ever rendered). Bumped quota_height to 5 (2 chrome + header + 5h + 7d). State transitions (tests/tui_state_test.rs, 20 tests): - toggle_sort cycles all 5 columns including Type wrap - toggle_time_filter cycles all 4 windows - select_next/previous wrap-around for both Tools and Live focus - toggle_focus auto-fallback when no live sessions exist - cycle_agent / cycle_project with empty + populated lists - refresh is no-op when paused Render assertions (tests/tui_render_test.rs, 11 tests): - Empty state hides the Live panel - CTX column shows 'used/window pct%' format - STATUS column color-codes (RateLimited literal renders) - Subagents inline into TASK column - Quota panel renders 5h + 7d gauges with percentages - Orphan ports strip surfaces port + command - Host vitals strip shows CPU/MEM/LOAD with values - Tools table TYPE column distinguishes builtin vs mcp - Footer reports current focus - Selected live session detail strip shows children + ports - Paused indicator in header
Extracted the pure orphan-detection logic from Scraper::detect_orphan_ports into a free function compute_orphans_and_evictions, parameterized on an is-alive closure. The Scraper method retains its self-mutating behavior; the helper enables testing without instantiating a real ProcessScanner. Tests added (scraper/mod.rs internal): - compute_orphans_finds_port_with_dead_parent - compute_orphans_evicts_dead_pids - compute_orphans_skips_live_sessions - rate_limit_info_is_at_limit_handles_boundaries (98.9% < 99% threshold) - rate_limit_only_seven_day_at_limit - session_status_label_round_trips
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Patch release that fixes a silent bug in the approval-rate column and brings the README's Limitations section in line with reality. Surfaced while verifying each README claim end-to-end against real DuckDB data.
Fixes
get_tool_metricswas querying$.decisionontool_resultevents, but Claude Code emitsdecision_typeinstead — every row was NULL and APR% silently fell back to 100% for every tool. Rewrote the query to:tool_result.decision_type(value:accept)tool_decision.decision(the only place rejects surface — rejected calls never produce atool_result)tool_use_idjoin, so they attribute to the propermcp__server__toolname instead of the genericmcp_toolthattool_decisionevents carryTaskCreate,TaskUpdate,TaskOutput,TaskStop,TaskList,TaskGet,ToolSearch,TestRead. They were rendering asmcpin the TYPE column.Tests (the missing verification layer)
test_approval_rate_from_real_event_shape— round-trips realistic Claude-Code event shapes throughduckdb-rs; asserts thetool_use_idjoin works and no phantommcp_toolrow leaks.test_app_approval_rate_end_to_end— fullApp → storage → tool_metricspipeline. ExitPlanMode 25% (1/4), Bash 67% (2/3).Docs
used/windowin the Live sessions panel.tool_use_id).Real-data verification
Querying live DuckDB after the fix produces actually useful numbers (was 100% across the board before):
Test plan
cargo test— 272 tests passing (added 2 round-trip tests)cargo clippy --all-targets -- -D warnings— cleancargo fmt --all -- --check— clean~/Library/Application Support/agenttop/metrics.duckdbviaduckdbCLI