Project: AgentFlow Document date: 2026-04-18 Repository snapshot reviewed: 2026-04-18 Updated: 2026-09-07 (failed-auth throttle: client identity taken from the right of X-Forwarded-For, admin window separated, a valid key always served; production chart must declare whose address the pod observes) Audience: engineering, security review, enterprise due diligence
AgentFlow exposes a public FastAPI surface for AI agents and tenant-owned integrations. The current repository shows a security posture centered on typed request validation, tenant-scoped access control, API-key authentication, rate limiting, SQL safety guards for NL-to-SQL, and CI-based dependency and image scanning. Response-side PII masking is no longer on this list: it was removed 2026-07-01 because the demo serving warehouse holds no PII, and real PII is governed engine-side in the DV2 vault (see section 7).
The codebase is strongest at application-layer controls that can be validated directly in source: auth, authorization, request filtering, security headers, contract evolution, replay safety, and auditability of API usage. The weakest areas are the controls that typically require external infrastructure or third-party attestation. In this repository snapshot there is no evidence of an external penetration test or demonstrated generalized secrets manager integration. Local DuckDB files can be opened through an optional encrypted attach path when an operator supplies encryption key material, but the default remains backward-compatible and unencrypted.
External pen-test attestation status as of 2026-05-06: not present. Use
docs/operations/third-party-pen-test-intake.md for the checklist required
before any third-party pen-test claim.
Automated posture channel (2026-06-05): the repository now publishes an
OpenSSF Scorecard result (.github/workflows/scorecard.yml) and carries a
prepared OpenSSF Best Practices self-assessment
(docs/operations/openssf-security-posture.md). Both are $0 posture signals —
an automated heuristic score and a maintainer self-certification respectively —
and are kept explicitly distinct from a third-party penetration test, which is
still not present and not claimed.
Threat model assumed by the current implementation:
- untrusted external callers using
X-API-Key - tenant isolation requirements across shared serving infrastructure
- AI agents issuing natural-language queries that must not escape allowed tables or mutate data
- operational abuse such as brute-force authentication attempts and burst traffic
The API uses tenant-bound API keys plus a separate admin secret for /v1/admin/*. New hashed API-key material uses Argon2id by default and is paired with a deterministic, peppered HMAC lookup digest so authentication performs one slow verification instead of scanning every key. Existing bcrypt hashes remain supported as a legacy migration path; bcrypt_rounds: 12 applies only when bcrypt is selected explicitly.
Rotation support is implemented in the auth layer. Keys have key_id, previous_key_hash, previous_key_active_until, and explicit grace-period behavior. Admin rotation endpoints expose create, rotate, rotation-status, and revoke-old flows. The auth middleware also records endpoint usage per tenant/key, which gives the system a concrete audit trail for key activity and key-slot transitions.
That lookup digest is only as private as its pepper, and the built-in one (agentflow-key-lookup-v1) is a constant in security.py. Left at the default, key_lookup is an HMAC anyone can recompute: a leaked api_keys.yaml then lets a guessed key be confirmed against the stored digest with one HMAC and no Argon2id verify, and lets two deployments' digests be joined into one identity. AGENTFLOW_PROFILE=production therefore refuses to boot when AGENTFLOW_KEY_LOOKUP_PEPPER is unset or equal to that constant (audit FB-07) — the same rule the query-analytics fingerprint pepper has had since AF-13, whose check moved to the same boot gate so a production pod can no longer come up and fail later on the analytics path. Both are supplied under Helm through extraEnv as valueFrom.secretKeyRef; templates/production-contract.yaml refuses a profile=production render that omits either or writes one as a literal value. Dev and demo keep the defaults so a fresh checkout still starts. Rotating the lookup pepper invalidates every stored key_lookup and drops those keys back to the O(n) verify scan until they are re-issued (docs/runbooks/auth-401-spike.md).
Authorization is layered on top of authentication:
- admin endpoints require
X-Admin-Key - entity access can be restricted per key through
allowed_entity_types - request context binds
tenant_idto the authenticated tenant - serving paths use that tenant context when querying tenant-scoped data
Evidence: config/security.yaml, src/agentflow_runtime/serving/api/security.py, src/agentflow_runtime/serving/api/auth/manager.py, src/agentflow_runtime/serving/api/auth/middleware.py, src/agentflow_runtime/serving/api/auth/key_rotation.py, helm/agentflow/templates/production-contract.yaml, tests/unit/test_auth_argon2_lookup.py, tests/unit/test_key_lookup_pepper_gate.py, tests/unit/test_helm_production_values_contract.py, tests/integration/test_rotation.py
/metrics is mounted without API-key auth so Prometheus can scrape it in-cluster through the default ClusterIP Service. The liveness and readiness probe paths are exempt for the same reason: kubelet dials the Pod IP and never traverses Ingress. /docs* and /openapi* are also API-key-exempt in _is_exempt_path, but production_docs_guard (src/agentflow_runtime/serving/api/main.py) answers 404 for /docs, /redoc and /openapi* whenever config.profile=production (audit G-2), so Swagger UI (/docs) and the schema (/openapi*) are reachable without a key on the dev/demo profiles only; /redoc is not in _is_exempt_path, so it needs an X-API-Key on the dev/demo profiles; on production production_docs_guard runs ahead of the auth middleware and answers 404 to every caller, key or not. The /docs, /redoc and /openapi.json entries in the production ingress prefix list therefore route to that 404. /v1/node/events authenticates with its own node bearer token (ADR 0012).
helm/agentflow/templates/production-contract.yaml refuses a production values file whose ingress rules would send /metrics to the API — a / Prefix path, /metrics under Prefix or Exact, a path that is not a canonical single-line absolute path (unquoted, a CR, LF or tab used to inject a second Ingress rule for /metrics), a className that is not a canonical single line (unquoted, it used to inject spec.defaultBackend and send unmatched requests, /metrics included, to the API), a host that is not a canonical single line (.host has always been quoted, but a multi-line host makes the rendered Ingress unprovable, so the contract refuses it defensively), or a pathType other than Prefix/Exact. It also refuses the exact ingress.annotations keys rewrite-target, use-regex, app-root, configuration-snippet, and server-snippet under both nginx.ingress.kubernetes.io/ and legacy ingress.kubernetes.io/: ingress-nginx interprets them after Helm checks the literal host/path, so they can change controller matching or the upstream URI. The path clauses are a denylist: a production host with an empty paths list satisfies them vacuously — the render then carries a rule with no paths, which routes nothing and is rejected on apply. helm/agentflow/templates/ingress.yaml quotes or toYaml-serialises every user-controlled ingress interpolation, so an injected value stays a scalar. Production also requires networkPolicy.enabled=true, and helm/agentflow/templates/networkpolicy.yaml limits pod ingress to the namespaces in networkPolicy.ingressFromNamespaces on the service port. /metrics stays unauthenticated by design for that in-cluster scrape. The chart default for networkPolicy.ingressFromNamespaces (helm/agentflow/values.yaml) enumerates only ingress-nginx. values-production.yaml keeps that same single selector and does not ship a guessed scrape namespace; the operator must name the namespace that runs Prometheus, or set networkPolicy.scrapeFromIngressNamespace=true when Prometheus shares the ingress-controller namespace. An empty ingressFromNamespaces list renders ingress: [] (deny all), so neither the ingress controller nor in-cluster Prometheus can reach the service port; the production contract refuses that empty list for that reason. A non-empty list whose every kubernetes.io/metadata.name selector is ingress-nginx (and no entry lacks that key — an entry selecting by another label counts as other and passes) is refused because it blocks scrape from a dedicated monitoring namespace, unless networkPolicy.scrapeFromIngressNamespace=true records the co-location as a deliberate decision. Production also requires service.type=ClusterIP: NodePort/LoadBalancer would publish the service port (/metrics included) without any Ingress rule, and ExternalName would void the routing contract.
This is a values-contract check at helm template time for the production profile only; it does not authenticate /metrics, and it does not bind the chart's deliberately dev-shaped defaults (a / path still renders green there). The annotation denylist covers only the Ingress object rendered by this chart; it cannot constrain the ingress-nginx controller ConfigMap or separately managed Ingress objects. The 2026-09-02 audit's F-10 acceptance criteria are therefore only partially met: no monitoring identity (mTLS, auth proxy, or IP allowlist) is implemented, and those cluster-level routing inputs remain operator-controlled. Production binds service.type to ClusterIP so NodePort/LoadBalancer cannot publish the service port (/metrics included) on a green render; ingress.enabled=false remains the sanctioned external-gateway shape and moves routing (and the /metrics exposure question) outside the chart.
Evidence: src/agentflow_runtime/serving/api/main.py, src/agentflow_runtime/serving/api/auth/middleware.py, helm/agentflow/templates/production-contract.yaml, helm/agentflow/templates/ingress.yaml, helm/agentflow/templates/networkpolicy.yaml, helm/agentflow/values-production.yaml, tests/unit/test_metrics_not_publicly_routed.py, tests/unit/test_helm_production_ingress_annotations.py, tests/unit/test_helm_production_scrape_networkpolicy.py, tests/unit/test_helm_production_service_type.py, tests/unit/test_docs_profile_gate.py
This section previously described a mechanism that did not work. It said the boundary was a schema qualification — TenantRouter mapping a tenant to a DuckDB schema name, and the SQL builder qualifying tables into "acme"."orders_v2". Nothing in src/ ever issued CREATE SCHEMA, so that relation did not exist at runtime and every authenticated entity read failed; on ClickHouse the same qualification named a database nobody creates. The suite was green because the shipped keys named a tenant absent from config/tenants.yaml, so the qualification resolved to nothing and silently did not apply. A boundary no test can tell apart from its own absence is not a boundary. See ADR-004.
The boundary is now the tenant_id column, on both stores, and it is part of each serving table's write key — ORDER BY (tenant_id, <pk>) on ClickHouse, PRIMARY KEY (tenant_id, <pk>) on DuckDB. That distinction is load-bearing: with a single-column key, two tenants' rows sharing an order_id are two versions of one ReplacingMergeTree row, and the later insert destroys the earlier. No read-side filter can undo that, so a filter alone was never sufficient.
Reads pass through one chokepoint. SQLBuilderMixin._qualify_table returns a tenant-filtered sub-select rather than a name, _scope_sql performs the same substitution inside metric templates and NL-generated SQL via the sqlglot AST, the search index carries the tenant on each document and filters before scoring, and the journal applies its own predicate. An unscoped read against a store that holds more than one tenant's rows is refused (503), not answered.
What the evidence supports today:
- DuckDB — proven. Two tenants with identical
order_idresolve to different rows; cross-tenant lookups return404; aggregates sum only the caller's rows; an unscoped read fails closed. Property tests assert the invariant over generated tenant and entity ids, not just the two-tenant example. - ClickHouse — the same model is implemented, provisioned, and proven live.
assert_tenant_key()refuses to serve a store still on the old sorting key;provision --migraterebuilds one. The adversarial two-tenant suite (tests/integration/test_clickhouse_tenant_isolation_live.py) runs against ClickHouse 25.3 on the CItest-integrationjob and was green on the audit stand (full live suite including tenant isolation). Multi-tenant ClickHouse is a supported claim for the write-key + read-chokepoint model described above.
It does not support broader claims such as end-to-end isolation across every external dependency.
Evidence: docs/decisions/004-tenant-id-column-over-schema-per-tenant.md, src/agentflow_runtime/tenancy.py, src/agentflow_runtime/serving/semantic_layer/query/sql_builder.py, src/agentflow_runtime/serving/backends/clickhouse_backend.py, tests/integration/test_tenant_isolation.py, tests/property/test_tenant_isolation_properties.py, tests/integration/test_clickhouse_tenant_isolation_live.py
Typed validation is pervasive in the API surface. FastAPI request bodies and query parameters are defined with Pydantic models across agent, batch, alert, webhook, dead-letter, SLO, and contract endpoints. Validation constraints are used for lengths, enums, numeric ranges, and optional structures.
The ingestion schemas add cross-field semantics beyond shape validation. OrderEvent verifies that total_amount matches the sum of line items, payment timestamps are normalized to UTC and rejected if too far in the future, and product pricing rejects negative values. This matters because it prevents upstream data corruption from turning into trusted downstream state.
Schema contract evolution is implemented through a contract registry plus version-aware validation and diff endpoints. The API versioning layer also exposes deprecation metadata through headers and supports tenant-level version pins, which reduces the blast radius of backward-incompatible changes.
Evidence: src/agentflow_runtime/ingestion/schemas/events.py, tests/unit/test_event_schemas.py, src/agentflow_runtime/serving/semantic_layer/contract_registry.py, src/agentflow_runtime/serving/api/routers/contracts.py, src/agentflow_runtime/serving/api/versioning.py
The serving layer uses two complementary patterns for SQL safety.
First, the hot-path entity and metric lookups pass untrusted values as query parameters rather than interpolating them into SQL text. Injection-focused unit tests assert that payloads such as '; DROP TABLE ... stay in parameter arrays and never appear in the generated SQL.
Second, the NL-to-SQL surface validates translated SQL with sqlglot. The validator only permits a single SELECT statement, rejects DDL and DML node types, and rejects unknown tables outside the allowlist and CTE names. Tenant scoping is then applied through AST-aware rewriting in _scope_sql, which is materially safer than regex replacement.
The repository also documents intentional # nosec B608 suppressions only on trusted identifier paths where identifiers come from internal catalog/config allowlists rather than user-controlled input.
Audit finding A-4 flagged the dynamic-SQL surface as "one careless edit away from a hole". Every # nosec B608 suppression in src/ was reviewed and is safe by construction:
- Interpolated identifiers come from a fixed in-code allowlist, the semantic catalog, trusted backend config, live schema introspection (
PRAGMA table_info),_IDENTIFIER_RE, orsqlglotvalidation. - Interpolated values are either bound as
?parameters, fixed literals, integers, or regex-extracted tokens that exclude SQL metacharacters and are additionally quoted via_quote_literal/_sql_str_literal.
No site interpolates unbound, unquoted request data. There are no Class-A (migratable value-interpolation) sites remaining: the hot entity/metric paths already bind values (use_query_params on DuckDB, _quote_literal elsewhere) and the operational routers already pass values through ?. The remaining suppressions are Class-B (identifiers / structural fragments that cannot be parameterized). The PostgresControlPlaneStore sites added by ADR 0010 slice 5 (control_plane/postgres.py, reviewed 2026-07-03) interpolate only a table name that is a module literal at exactly two call sites; every value binds via %s.
audit P0-3 (2026-07-11). The five journal sites moved out of routers/lineage.py (1) and routers/slo.py (4) into semantic_layer/journal.py, which reads pipeline_events through the active serving backend instead of a private DuckDB cursor. The interpolated surface is unchanged in kind, and smaller in spread: every fragment is an identifier taken from a live schema probe (table_columns) against a fixed allowlist — the time column, the nullable-column fallbacks — plus the SLO quantile, a float from config/slo.yaml. Values still never interpolate: JournalReader._value() binds them as ? on DuckDB and _quote_literal-escapes them on ClickHouse, whose execute(params=...) is a documented no-op. So entity_id, which arrives in the URL path, binds exactly as it did before the move. The one exception is the time window ('7 days'), inlined on both backends because INTERVAL ? is a DuckDB syntax error and CAST(? AS INTERVAL) has no ClickHouse translation — it is derived from an int in config/slo.yaml and is never request data. entity_queries.py gained the bulk entity scan the search index used to run on the raw connection (catalog-defined table name, int() limit, no values).
The number of suppressions per file is pinned by test_interpolated_sql_nosec_surface_is_pinned — a new site (even inside an already-listed file) or a new file fails CI and forces a review. Each suppression's per-line rationale comment is enforced by test_nosec_comments_carry_reason.
| Site | Interpolated | Why safe (Class B — safe by construction) |
|---|---|---|
semantic_layer/nl_engine.py:110,163 |
oid, uid (values) |
regex ORD-[\w-]+ / USR-\d+ exclude quotes; additionally quoted via _sql_str_literal; covered by test_nl_engine_injection.py |
semantic_layer/nl_engine.py:116,126,149 |
window (value) |
_extract_window numeric allowlist (\d+ <unit> or constant); quoted via _sql_str_literal |
semantic_layer/nl_engine.py:139 |
limit (value) |
parsed int() from the question text |
api/routers/lineage.py:103 |
select_columns, time_column (identifiers) |
column names are in-code literals gated by PRAGMA table_info; values bound as ? |
api/routers/slo.py:109,123,143,160 |
time_column (identifier) |
fixed allowlist processed_at/created_at; window/tenant bound via CAST(? AS INTERVAL) / ? |
api/routers/stream.py:45 |
select_columns (identifiers) |
in-code literals gated by PRAGMA table_info; filters bound as ? |
api/webhook_dispatcher.py:315 |
order_by (identifier) |
fixed allowlist processed_at/created_at/event_id; tenant bound as ? |
backends/clickhouse_backend.py:293,306,326,344,359,375 |
self._database + fixed table names |
database from trusted backend config; table names and VALUES are in-code literals (demo seed) |
backends/duckdb_backend.py:94 |
table_name (identifier) |
guarded by _IDENTIFIER_RE.match (bare identifier or schema.identifier) |
backends/duckdb_backend.py:116 |
sql (full statement) |
sqlglot.parse must yield exactly one exp.Select before EXPLAIN |
semantic_layer/query/entity_queries.py:31,133,196 |
table / primary_key / entity_type | identifiers via _quote_identifier from catalog; values via _quote_literal or ? (use_query_params on DuckDB) |
semantic_layer/query/nl_queries.py:133,138 |
sql subquery, limit/offset |
sql prevalidated by validate_nl_sql (sqlglot single SELECT); limit bounded 1..1000, offset decoded int |
semantic_layer/search_index.py:152 |
entity.table (identifier) |
table name from the semantic catalog EntityDefinition, not request data |
orchestration/dags/daily_batch.py:148 |
table (identifier) |
iterates a fixed in-code health-check list |
Evidence: src/agentflow_runtime/serving/semantic_layer/sql_guard.py, src/agentflow_runtime/serving/semantic_layer/query/sql_builder.py, src/agentflow_runtime/serving/semantic_layer/nl_engine.py, tests/unit/test_query_engine_injection.py, tests/unit/test_sql_guard.py, tests/unit/test_nl_engine_injection.py, tests/unit/test_security_tooling_policy.py
Rate limiting exists at two levels:
- per-key request quotas enforced through
RateLimiter - per-IP throttling for repeated failed authentication attempts
RateLimiter uses Redis when available and falls back to an in-memory sliding window when Redis is unavailable or intentionally disabled. The auth middleware returns X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset, and Retry-After headers, so clients can adapt to the policy rather than blindly retrying.
The threshold is security.max_failed_auth_per_ip_per_hour (default 10, canonical config/security.yaml) over a one-hour sliding window held per process. It exists to blunt scanning and log-flooding, not to stop brute force: API keys are 256-bit random values, so guessing was never the thing it was holding back. Three properties follow from that, and each of them is a deliberate limit rather than an oversight (audit FB-06):
- A valid key is always served. The middleware resolves the presented key before it consults the window; only a failed attempt is counted, and only a failed attempt is answered with 429. The earlier order — reject on the window, then look at the key — meant eleven requests with any junk key took every caller sharing that address offline for an hour. Under throttle the resolution is capped to the constant-work paths (runtime cache and the O(1) peppered lookup), so a key issued since M-C4 still authenticates for exactly one hash while a scanner cannot buy N bcrypt verifications per guess. A pre-M-C4 bcrypt entry with no
key_lookupis not resolvable while its address is throttled; rotate it onto an argon2id entry. - The admin surface has its own window. Failures on
X-Admin-Keyare counted separately from failures onX-API-Key, so a scan against/v1cannot throttle/v1/admin— the surface an operator needs while the scan is running — and vice versa. - The client address is only as truthful as the deployment makes it.
X-Forwarded-Foris honoured solely when the immediate peer is listed inAGENTFLOW_TRUSTED_PROXIES, and the chain is then read right to left, skipping hops that are themselves trusted proxies and stopping at the first hop no proxy vouches for. The leftmost element is client-supplied and is never used. With no trusted proxies configured behind a gateway, every caller shares one address: the throttle is then best-effort — it will not lock anyone out, but any legitimate request from that address clears the window. Naming the proxies (or declaringconfig.gateway.preservesClientIp) is what makes it meaningful, andprofile=productionrefuses a release that answers neither.
The window is per process by design: on N replicas an attacker spreading guesses gets N times the budget, which is accepted while key entropy is what it is. It is bounded, too — once a window is over the limit it stops accumulating, so a scanner cannot grow it by one entry per request.
The SDKs also include resilience primitives, specifically retry policy handling and a circuit breaker. This is not a server-side abuse control by itself, but it does reduce retry storms and repeated hammering of degraded endpoints by well-behaved clients.
Evidence: src/agentflow_runtime/serving/api/rate_limiter.py, src/agentflow_runtime/serving/api/auth/middleware.py, sdk/agentflow/client.py, sdk-ts/src/client.ts, tests/unit/test_sdk_circuit_breaker.py, tests/unit/test_auth_throttle_identity.py
Superseded (2026-07-01). The response-side PII masking / deny-gate described in this section has been removed. The demo serving warehouse holds no PII (
users_enriched/orders_v2carry only analytics columns), so the machinery guarded an empty surface. Real contact PII lives only in the DV2 business vault and is governed engine-side there (ClickHouse row/column policies — ADR 0006 Phase 2). The paragraph below is retained as the point-in-time record.
Response-side PII masking was implemented in a PiiMasker module and applied on entity responses and NL-query results. Masking behavior was configured through a pii_fields config file, supported multiple strategies (partial, full, hash), and allowed explicit tenant exemptions for internal tenants. When masking was applied, the API set X-PII-Masked: true. The module, its config, and its tests are gone from the tree; the removal is recorded in the CHANGELOG (2026-07-01) and the engine-side replacement in ADR 0006 Phase 2.
Security headers are applied centrally and include:
Strict-Transport-SecurityContent-Security-PolicyX-Frame-OptionsX-Content-Type-OptionsReferrer-Policy
These controls improve baseline browser-facing hardening for docs/admin surfaces. TLS termination is intentionally delegated to an upstream edge or ingress layer; the FastAPI application applies HTTP-layer security controls behind that boundary.
Evidence: src/agentflow_runtime/serving/api/security.py (security headers, still current); the removed masker's history lives in the CHANGELOG (2026-07-01) and docs/decisions/0006-fix-demo-serving-engine-on-clickhouse.md Phase 2.
The repository includes a dedicated security workflow in GitHub Actions:
- Bandit for Python static analysis with a tracked baseline diff
- Safety for dependency vulnerability scanning
- Trivy for container image scanning and CycloneDX SBOM generation
The security workflow now scans both shipping runtime images:
agentflow-api:security-scan and agentflow-flink:security-scan.
Each image produces distinct CycloneDX, JSON, and SARIF outputs.
JSON and SARIF scans use HIGH,CRITICAL plus ignore-unfixed; report
generation uses exit-code 0 so the waiver-aware evaluator, rather
than the raw scanner, decides policy. scripts/evaluate_trivy_policy.py
runs once for api-runtime and once for flink-runtime against
security/trivy-waivers.json and fails for unwaived findings, expired
waivers, or stale waivers. API and Flink SARIF uploads have distinct
categories; SBOM artifacts also have distinct names. make trivy-policy
evaluates already-generated JSON reports for both scopes; it does not
build or scan images and does not suppress failures. JSON, SARIF, SBOM,
policy-summary, and IaC working files live under ignored
.artifacts/trivy/. They are replaceable runtime/CI artifacts, not
reviewed evidence or production acceptance; promotion requires a new
date-stamped identity with provenance. The api-runtime
policy scope is empty, so every API finding is unwaived; the existing
Flink scope retains its narrow expiring waivers.
Dependency-scan working files follow the same ownership. The Bandit job
writes its raw JSON report to ignored .artifacts/security/bandit-current.json
and scripts/bandit_diff.py compares it against the tracked
.bandit-baseline.json, which is the only reviewed input. The Safety job
resolves its requirement buckets, resolver virtualenvs, and the
vulnerable-pin regression probe under .artifacts/security/safety/.
scripts/run_safety_scan.py runs safety check once per bucket and applies
--ignore only for waiver safety_id values whose scope matches that
bucket. The pip-audit job exports the full locked profile set to
.artifacts/security/pip-audit/; the production pip-audit step reads the
tracked requirements-docker.lock directly. These are replaceable per-run
working files, not reviewed evidence or production acceptance; promotion
requires a new date-stamped identity with source SHA, workflow run, scanner
versions, exact command/configuration, outcome, and hash provenance.
A version bump closes most findings. PYSEC-2026-3740 in nltk 3.10.3 closes
nothing: the advisory itself says "Patched versions: Not yet patched", so there
is no release to move to. Until 2026-09-07 the repository had nowhere to put
that. validate_waiver required a fixed_version, which meant an unfixed
finding was unwaivable by construction — its key ends in an empty fix version
and no valid waiver key could — and the pip-audit job had no waiver mechanism
at all, so it simply stayed red.
Three changes, none of which loosen the gate:
fixed_version: nullis now a statable claim — "upstream has published no fix" — and the key must still be present, so silence stays a typo rather than a claim.scripts/run_pip_audit_scan.pyruns pip-audit with no--ignore-vulnand evaluates the JSON report againstsecurity/trivy-waivers.jsonitself, scopepython-profiles. Unwaived findings, expired waivers, and waivers that match nothing each fail the job, exactly as the Trivy path does.- The fix state is part of the match. The nltk waiver's premise is that no fix exists; the day pip-audit reports one, the waiver stops matching, the finding returns to unwaived, and the gate goes red on the release that is now available. Nobody has to remember to revisit it.
The argument for this particular waiver is narrow. The affected APIs are nltk's
model-persistence helpers, and the bypass only matters to a caller that enables
nltk's pathsec sandbox and lets untrusted input choose model paths. AgentFlow
does neither: nltk arrives only as a transitive dependency of llama-index-core
under the integrations extra, no source file references nltk or pathsec
(a test asserts this), and the package is absent from requirements-docker.lock,
so it does not ship in the API image. The waiver expires 2026-11-01.
requirements-docker.lock is deliberately excluded from this mechanism. The one
inventory that ships is audited bare, with no waiver path reachable at all.
On 2026-07-30, a Trivy scan of the API image identified vulnerable packages
vendored by runtime pip, not dependencies from the application lock. The
final stage now removes pip, setuptools, and wheel after the hash-locked
install and pip check. An independent Mac rebuild and Trivy 0.70.0 scan
reported zero HIGH/CRITICAL findings while the API import remained healthy.
See
security-runtime-image-trivy-2026-07-30.md.
The same closeout also converted two implicit dependency assumptions into tested supply-chain boundaries: MCP is constrained to its supported 1.x major API, and PyIceberg's write-time native core is explicitly locked to the Python 3.11–3.13-compatible 0.7 line. The regenerated hash lock, clean Mac environment, rebuilt image, OSV queries, and Trivy result are recorded in dependency-compatibility-2026-07-30.md.
The Bandit baseline currently records a historical B310 finding in src/agentflow_runtime/serving/backends/clickhouse_backend.py. SQL construction findings are not globally suppressed; reviewed identifier construction is handled through narrow suppressions and tests.
Helm defaults no longer embed production-shaped API-key verifier hashes. Operators can render a chart-managed Secret for local use or mount an existing Kubernetes Secret, which is friendlier to External Secrets Operator, Sealed Secrets, or equivalent workflows.
Evidence: .github/workflows/security.yml, .bandit, .bandit-baseline.json,
docs/operations/helm-deployment.md,
docs/evidence/security-runtime-image-trivy-2026-07-30.md,
scripts/evaluate_trivy_policy.py, security/trivy-waivers.json,
scripts/run_pip_audit_scan.py, Makefile,
tests/unit/test_security_image_scan_policy.py,
tests/unit/test_security_workflow.py,
tests/unit/test_run_pip_audit_scan.py
Operationally, the repo shows several useful security-facing controls:
- API usage is written to
api_usagewith tenant, key name, endpoint, key ID, and key slot - admin analytics endpoints can inspect usage, anomalies, latency, and top entities/queries
- the runbook includes response procedures for API unavailability, pipeline lag, dead letters, webhook failures, alert storms, and stuck key rotation
This provides a credible audit and incident-response starting point for a small team. It is notably better than a pure demo API with no usage telemetry.
Refusals on the admin surface leave a record and not only a counter (audit FB-10). require_admin_key guards the routes that issue, rotate and revoke every tenant API key, and each of its three refusals — admin_invalid, rate_limited and admin_unconfigured — increments agentflow_auth_failures_total and emits a structured admin_auth_failed line carrying reason, client_ip, path and the redacted request headers. It is a separate event from the tenant path's api_auth_failed so that a scan against /v1 and someone guessing the operator credential stay distinguishable at query time. The line never carries key material: headers pass through security.sensitive_headers_to_redact and then lose X-Admin-Key unconditionally, because that list is operator-configurable and audit F-11 had already found a built-in default that omitted it. The label vocabulary in docs/runbooks/auth-401-spike.md § Detection is pinned against the labels the code actually emits, in both directions — it had been carrying one name nothing emitted and neither admin name.
The admin key itself is one shared value with no dual-key window: a pod validates against the single value it resolved at startup, so rotation is a Secret change plus a rolling restart during which admin calls are unreliable. docs/operations/admin-key-rotation.md owns that procedure — trigger conditions, the ordering that keeps the analytics-retention CronJob from failing mid-run, and how to confirm the old value is dead.
What is not evidenced in this repository snapshot:
- generalized secrets management through AWS Secrets Manager or another external vault
- automated rotation for non-API-key secrets
- externally immutable audit retention or SIEM export
The API usage path can optionally publish hash-chained JSONL records through AGENTFLOW_AUDIT_LOG_PATH in addition to DuckDB analytics. This is useful local evidence that DuckDB analytics are not the only audit path, but object-lock retention, SIEM delivery, and external immutability still need operator evidence outside the repository. Collect that operator evidence outside this repository before making any external immutable-retention claim.
Because the external controls are not provable from the checked-in code, they should not be claimed in customer-facing security questionnaires without additional infrastructure evidence.
Evidence: src/agentflow_runtime/serving/api/auth/middleware.py, src/agentflow_runtime/serving/api/analytics.py, docs/runbook.md, docs/runbooks/auth-401-spike.md, docs/operations/admin-key-rotation.md, tests/unit/test_admin_auth_audit_log.py
The current implementation has several material limitations:
- No external penetration test evidence is present in the repository.
- DuckDB encryption is optional and operator-configured; the default file path remains unencrypted for backward compatibility, and DuckDB encryption is not a compliance attestation by itself.
- Secrets management appears environment- and chart-driven in this snapshot; a managed secret store is not demonstrated.
- The security pipeline is strong at code scanning and SBOM generation, but a real signed container-release run is still evidence-pending until CI signs a published image digest.
- The demo-data initialization path is convenient for development, but it increases the importance of strict environment separation between demo and production deployments.
- Browser-oriented security headers exist, and request body size enforcement is applied from
SecurityPolicy.request_size_limit_bytes. /metricsremains unauthenticated to anything that can already reach the pod port; no monitoring identity (mTLS, auth proxy, or IP allowlist) is implemented. Production Ingress rules that would send/metricsto the API are refused athelm templatetime, as are the exact ingress-nginx routing-control annotationsrewrite-target,use-regex,app-root,configuration-snippet, andserver-snippetunder the current and legacy prefixes. That values-contract check does not authenticate the endpoint or bind the chart's dev/demo defaults, and its annotation denylist covers only the Ingress object rendered by this chart — not the ingress-nginx controller ConfigMap or separately managed Ingress objects. Production bindsservice.typetoClusterIP;ingress.enabled=falseremains the sanctioned external-gateway shape and moves routing (and the/metricsexposure question) outside the chart. Production NetworkPolicy must name a scrape namespace innetworkPolicy.ingressFromNamespaces(values-production.yamlships no guessed namespace; the unedited overlay is refused). An empty list rendersingress: [](deny all) and is refused because neither the ingress controller nor Prometheus can reach the service port. A list whose everykubernetes.io/metadata.nameisingress-nginxis refused unlessnetworkPolicy.scrapeFromIngressNamespace=truerecords that Prometheus shares the ingress-controller namespace. Egress is an allow-list only where a rule names destinations: a rule withports:and noto:permits that port to every address, in the cluster and on the internet, which is how the baseline came to allow 6379/9092/8181/9000/8123/4317/5432 anywhere (audit FB-09). Each rule now takes peers fromnetworkPolicy.egressTo.<service>, renders only when its feature is configured, and production refuses an empty peer list for every rule that renders. DNS remains the chart-owned exception, selecting kube-dns by label.- The admin key is a single shared credential. Every operator presents the same
X-Admin-Key, so an admin action is attributable to "someone holding the key" and nothing finer, and revoking one person's access means rotating it for everyone. There is no dual-key window during that rotation. Per-operator admin credentials, hashed the way tenant keys already are, would remove both properties and are not implemented;docs/operations/admin-key-rotation.mdis what stands in for them today.
Partial readiness. The repo supports tenant scoping, usage auditing, and engine-side PII governance in the DV2 vault (fail-closed column grants, row policies — ADR 0006 Phase 2). That helps with least-privilege data exposure and auditability. However, GDPR readiness is incomplete without documented data retention policies, deletion workflows, subject access procedures, and infrastructure evidence for storage/backup handling.
Partial readiness. Access control, audit logging, CI security scans, and operational runbooks are present. Missing evidence includes change-management controls outside git/CI, vendor management, centralized secrets management, formal incident program artifacts, and third-party audit evidence.
Not ready based on the reviewed repository. Optional DuckDB at-rest encryption does not supply the broader administrative, contractual, audit-retention, or external-assessment controls HIPAA would require.
For an engineering-led v1 product, AgentFlow shows an above-average application security baseline:
- strong typed validation
- practical tenant isolation
- real SQL safety controls
- concrete key rotation mechanics
- engine-side PII governance in the DV2 vault (the serving tier holds no PII)
- usable audit trail and CI scanning
The main gap is not the app layer. It is the absence of externally verifiable infrastructure and governance controls. These gaps do not block continued development, package publication, or demos, but they do block enterprise-facing security claims:
- external penetration testing, including report scope, dates, severity summary, remediation map, retest status, and owner
- documented secrets-management architecture
- explicit encryption-at-rest posture for deployment targets
- external immutable retention evidence before claiming WORM/Object Lock/SIEM-backed audit logs
- a short customer-facing security overview aligned to the facts above