Skip to content

test: ratcheting API call-count tests for hot commands (catch N+1 regressions) #802

Description

@padak

Problem

kbagent has no guard against performance regressions in the number of HTTP calls a command makes. A change that adds one extra API request per item (N+1), or an accidental duplicate fetch, passes all tests and ships silently. For AI agents that call kbagent repeatedly, extra round-trips are the dominant latency cost (and they also add load on Keboola APIs).

Proposal

Add API call count tests: for the most frequently used commands, run the command against mocked HTTP (pytest-httpx, already used in ~35 test files) and assert the exact number (and ideally the sequence of method + path) of requests made.

  • Deterministic: request count does not depend on network or machine speed, so tests are not flaky.
  • Ratcheting: the expected count is committed; a PR that increases it must update the number explicitly (visible in review), a PR that reduces it lowers the number.
  • Parametrize over input size where it matters (e.g. 1 vs 10 items) so N+1 patterns are caught: the count should grow only where the design intends it (e.g. token list --with-last-used is documented as one extra call per token).

Suggested first scope

Start with the hot read paths agents use most, e.g.:

  • project list / project status
  • config list, config detail
  • job list (single project and multi-project fan-out), job detail
  • storage tables, storage table-detail
  • flow list, flow detail

Put a small shared helper in tests/helpers.py (e.g. a fixture that records requests and asserts against an expected list) so adding a new command to the suite is a few lines.

Acceptance criteria

  • Shared helper/fixture for asserting request count and sequence.
  • Call-count tests for the initial command set above, including at least one size-parametrized N+1 check.
  • Short note in CONTRIBUTING.md that new/changed commands on hot paths should add or update a call-count test.
  • Any existing redundant calls found while writing the tests are reported as follow-up issues (or fixed in the same PR if trivial).

Inspired by the "deterministic metrics + ratcheting CI guardrails" approach from https://claude.dev/blog/how-we-made-claude-ai-faster/

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions