Conversation
Users of the L2 announce table indexes only need to construct queries. They shouldn't depend on the statedb.Index values. Hence, export the L2AnnounceBy* helpers and keep the actual index definitions private to reduce the package API surface and to keep the index values internal. Signed-off-by: Tobias Klauser <tobias@cilium.io>
Users of the node addresses table indexes only need to construct queries. They shouldn't depend on the statedb.Index values. Hence, export the NodeAddressBy* helpers and keep the actual index definitions private to reduce the package API surface and to keep the index values internal. Signed-off-by: Tobias Klauser <tobias@cilium.io>
Users of the sysctl table index only need to construct queries. They shouldn't depend on the statedb.Index value. Hence, export the SysctlByName helper and keep the actual index definition private to reduce the package API surface and to keep the index value internal. Signed-off-by: Tobias Klauser <tobias@cilium.io>
Users of the ipset table index only need to construct queries. They shouldn't depend on the statedb.Index value. Hence, export the IPSetEntryByKey helper and keep the actual index definition private to reduce the package API surface and to keep the index value internal. Signed-off-by: Tobias Klauser <tobias@cilium.io>
The index is not queried outside the package. Unexport it to reduce the API surface and to keep the index values internal. Signed-off-by: Tobias Klauser <tobias@cilium.io>
Users of the BGP reconcile errors table indexes only need to construct queries. They shouldn't depend on the statedb.Index values. Hence, export the BGPReconcileErrorBy* helpers and keep the actual index definitions private to reduce the package API surface and to keep the index values internal. Signed-off-by: Tobias Klauser <tobias@cilium.io>
Users of the EDT table index only need to construct queries. They shouldn't depend on the statedb.Index value. Hence, export the EDTByID helper and keep the actual index definition private to reduce the package API surface and to keep the index value internal. Signed-off-by: Tobias Klauser <tobias@cilium.io>
Users of the enrolled namespaces table index only need to construct queries. They shouldn't depend on the statedb.Index value. Hence, export the EnrolledNamespaceByName helper and keep the actual index definition private to reduce the package API surface and to keep the index value internal. Signed-off-by: Tobias Klauser <tobias@cilium.io>
Users of the health status table indexes only need to construct queries. They shouldn't depend on the statedb.Index values. Hence, export the StatusByID and StatusByLevel helpers and keep the actual index definitions private to reduce the package API surface and to keep the index values internal. Now that the indexes are unexported, tests in other packaged can no longer easily construct the table, so export the constructor as NewTable. Signed-off-by: Tobias Klauser <tobias@cilium.io>
Remove the Gateway reconciliation warning for the LB IPAM IP annotation alias. The warning only covered a single annotation alias and could not reliably detect other infrastructure annotations that may overlap with base Gateway API functionality exposed by the CRD. It also logged from the Cilium operator, which surfaces the message to the operator rather than the Gateway user who would need to act on it. IMO there's also not much sense of removing the annotation from the list for similar reasons (as it was intended in the PR that introduced this warning log). We should simply treat this as user config issue. Signed-off-by: Marco Hofstetter <marco.hofstetter@isovalent.com>
The shared client machinery is concurrent, complex and a bit brittle. One particularly tricky aspect is that there's a channel which is being written to by some goroutines and closed by another, seemingly without proper synchronisation. However, the synchronisation lies in the fact that the upstream DNS server of miekg/dns does _not_ spawn goroutines for each query when they come in via TCP, it keeps one per connection. Fixing the channel ownership or synchronizing the goroutines each also have downsides, they make the concurrency even more pronounced and difficult to think about, hence I'm opting for a slightly unorthodox solution: "document" the assumption in form of a test which fails if the TCP queries are executed in parallel. The test is written with synctest, hence both fast and deterministic, as far as I understand. I've tested that it fails when I simply let each TCP handler run in its own goroutine (which would break other things, too), so it's sensitive enough. Reported-by: Mike Molchanov <mikier@google.com> Signed-off-by: David Bimmler <david.bimmler@isovalent.com>
Signed-off-by: Bohdan Krasko <bkrasko@cisco.com>
Signed-off-by: Bohdan Krasko <bkrasko@cisco.com>
Signed-off-by: Bohdan Krasko <bkrasko@cisco.com>
Replace hramos/needs-attention with a pinned actions/github-script step. The script preserves the existing behavior: when an issue author comments on an issue labeled need-more-info, remove that label and add info-completed. Also configure the image workflow's GitHub App token action with the app's client ID and corresponding AUTO_COMMENT_BOT_CLIENT_ID secret. Both workflows only run from main, so stable branches do not need this change. Signed-off-by: Bohdan Krasko <bkrasko@cisco.com>
These type GC fields are only written but never read. The respective values are already used to configure GC.heartbeatStore and GC.rateLimiter, respectively. There is no point in keeping these copies, so remove them. Signed-off-by: Tobias Klauser <tobias@cilium.io>
Replace the IPV4_SNAT_EXCLUSION_DST_CIDR* macros with a node-level runtime configuration. Populate the prefix from the native-routing CIDR, or from the configured IPv4 native-routing CIDR when the IP masquerade agent is enabled. Move union v4addr into ipv4_core.h so it can be used by node configuration structures. Signed-off-by: viktor-kurchenko <viktor.kurchenko@isovalent.com>
Fixup for c13cbcf ("bpf, datapath: move IPv4 direct routing address to runtime config"). Signed-off-by: Julian Wiedmann <jwi@isovalent.com>
Stop logging reconcile errors directly in the Gateway reconciler when the error is returned to controller-runtime, which already logs failed reconciliations. Wrap returned errors with step-specific context so the single emitted log still identifies the failing operation. Signed-off-by: Marco Hofstetter <marco.hofstetter@isovalent.com>
Stop logging reconcile errors directly in the GAMMA reconciler when the error is returned to controller-runtime, which already logs failed reconciliations. Wrap returned errors with step-specific context so the single emitted log still identifies the failing operation. Signed-off-by: Marco Hofstetter <marco.hofstetter@isovalent.com>
A newly created endpoint is exposed to the endpoint manager before its first regeneration completes. This allows endpoint initialization to overlap with a policy import. For identities unaffected by an import, UpdatePolicy normally advances the endpoint's realized revision without regenerating it. An initializing endpoint is handled differently because its policyRevision is still zero: UpdatePolicy returns without advancing it, assuming that the already queued initial regeneration will realize the latest revision. That assumption is not always true. Policy computation for the endpoint may already have selected the previous repository revision before the import occurs. Since the import does not affect this endpoint's identity, the policy computer does not schedule another computation for it. The initial regeneration therefore completes with the previous revision, and the endpoint remains behind until another policy update or periodic regeneration occurs. This was observed happening with a new endpoint in the CI IPsec connectivity job. Its identity policy was selected at revision N immediately before an unrelated policy advanced the repository to revision N+1. The initial and periodic regenerations both completed at revision N, causing the connectivity test's policy wait to time out. When UpdatePolicy encounters an unaffected endpoint whose realized revision is still zero, explicitly schedule policy computation for its identity at the new revision. Preserve that revision in skippedPolicyRevision so that the queued initial regeneration waits for the computation before proceeding. UpdatePolicy and endpoint regeneration are serialized by buildMutex. If regeneration has already started, UpdatePolicy waits for it and follows the normal nonzero-revision path. Otherwise, the queued regeneration consumes the deferred revision after UpdatePolicy releases the lock. Add a unit test covering the initializing, unaffected endpoint case. The test verifies that policy computation is scheduled at the target revision, the revision is not reported as realized prematurely, and the queued regeneration consumes and clears the deferred target. Signed-off-by: Jarno Rajahalme <jarno@isovalent.com>
Protect the test-only mockMetrics helper with a mutex so concurrent ACK, NACK, and cancel updates do not write to shared maps unsafely. This fixes a panic in xDS tests where multiple request streams can share the same mockMetrics instance and trigger "fatal error: concurrent map writes" while updating counters. Signed-off-by: Marco Hofstetter <marco.hofstetter@isovalent.com>
This is how original DSR ICMP code (removed in 90c05a6 ("bpf: dsr: fix outer SrcIP in ICMP error msg")) added its free space for the outer headers. Signed-off-by: Julian Wiedmann <jwi@isovalent.com>
Rather than having such a stand-alone length variable sitting around, connect the various header lengths right into the bounds check. Signed-off-by: Julian Wiedmann <jwi@isovalent.com>
Avoid copying out various bits of the inner L3 header. Signed-off-by: Julian Wiedmann <jwi@isovalent.com>
…vate Stop embedding controller-runtime's client.Client in the gateway and GAMMA reconcilers. Store it as an explicit private field instead, and unexport the scheme field as well. This keeps the reconciler surface narrower, makes dependencies more explicit, avoids promoting the full client interface through embedding, and removes embedding-related warnings. Update the gateway-api reconciler setup, reconcile paths, and in-package tests to use the new field names. Signed-off-by: Marco Hofstetter <marco.hofstetter@isovalent.com>
Signed-off-by: weizhoublue <weizhou.lan@daocloud.io>
The cilium-nodes-watcher one-shot job bootstraps the cloud IPAM
allocator, including the blocking initial synchronization with the
cloud API. Before the switch to modular IPAM allocators, a failure
there was fatal: `operator/cmd/root.go` called `logging.Fatal("Unable to
start %s allocator")` and the operator exited, letting Kubernetes
restart it.
Moving that bootstrap into a `job.OneShot` dropped the fatality, as the
job is registered without `job.WithShutdown()`. An unrecoverable failure
is now only logged at error level by the job runner and marks the job
degraded, leaving the operator running with no allocator and no node
watcher: it can no longer hand out any IP, yet stays alive with no
signal beyond a degraded health report.
Add `job.WithShutdown()` to restore the previous behavior. The hive is
shut down with the error, which `operator/cmd/root.go` turns into a
single fatal log. Combined with the error chain repaired in
f54b082 ("ipam: Return error from instance (re)sync instead of
sentinel time.Time"), that log now also carries the underlying cloud
API error, which the pre-regression fatal used to swallow.
The `job.OneShot` runner returns early when the job succeeds or when the
context is canceled, so a normal shutdown does not trigger this path.
The node watcher is shared with the clusterpool and multipool
allocators, which are equally unable to operate if their watcher fails
to start, so they get the same treatment.
Fixes: 9664b9b ("operator/ipam: Switch to modular IPAM allocators")
Signed-off-by: Hadrien Patte <hadrien.patte@datadoghq.com>
Move the cluster nodes REST API into its own cell and serve snapshots directly from Table[*node.Node]. Replace per-client NodeHandler subscriptions with StateDB change iterators. Track each client's last observed nodes to preserve add, update, delete and coalescing semantics while ignoring reconciler status-only updates. This removes the REST API's NodeManager dependency while retaining the legacy GetNodes API for external consumers. AIL:3 Related: cilium#41744 Signed-off-by: Jussi Maki <jussi@isovalent.com>
golangci-lint is compiled with the Go toolchain a branch installs, so it must not be updated on a stable branch without the matching Golang update. Cap it per branch in the Renovate configuration, and group the entries by branch rather than by dependency, so each Go version is one bucket holding both ceilings and raising one is visibly raising the other. The Golang update instructions in dev_setup.rst describe the bucket and say to raise both ceilings in the same change. This commit was prepared with AIL:3. I personally checked the docs-builder check-build.sh run over the rendered documentation. Signed-off-by: André Martins <andre@cilium.io>
BPF Host Routing is causing a number of issues when combined with Istio, Kind, or for various other setups that rely on iptables. Because it's transparently enabled once BPF masquerading and KPR are enabled, users often incorrectly blame BPF masquerading instead of BPF Host Routing. Let's disable it by default to avoid breaking users' setups and to be able to enable BPF masquerading by default. For Helm deployments with the upgradeCompatibility flag, the existing value will be preserved. Since it's now only enabled explicitly by users, we need to change the Info log to a Warning log when falling back to legacy host routing. Signed-off-by: Paul Chaignon <paul.chaignon@gmail.com>
Enables BPF masquerading by default on new Helm deployments of Cilium. Signed-off-by: Paul Chaignon <paul.chaignon@gmail.com>
Once BPF masquerading is enabled, the agent figures that devices are
required (cf. AreDevicesRequired()) and fails if a direct routing device
isn't found. Unfortunately, it cannot detect a direct routing device in
the case of the Endpoint integration tests (endpoint_test.go) because
the node IP addresses aren't set up and DirectRoutingDevice.Get()
therefore fails to find a device that matches the node k8s IP address:
unable to determine direct routing device. Use --direct-routing-device to specify it
Let's instead set the Option.DirectRoutingDevice to manually select
the direct routing device.
Signed-off-by: Paul Chaignon <paul.chaignon@gmail.com>
Sometimes, the workflow fails because it's running out of disk space
while building Cilium to check the generated documentation:
CGO_ENABLED=0 GOARCH=amd64 go build -mod=vendor -ldflags '-X "github.com/cilium/cilium/pkg/version.ciliumVersion=1.21.0-dev 939f936c106 2026-09-03T06:51:51Z" -s -w -X "github.com/cilium/cilium/pkg/envoy.requiredEnvoyVersionSHA=0512632eff02ed4971b16926ae6b82a3c0e07a70" ' -tags=osusergo,ipam_provider_aws,ipam_provider_azure,ipam_provider_operator,ipam_provider_alibabacloud -o cilium-operator
github.com/cilium/cilium/operator: go build github.com/cilium/cilium/operator: copying /tmp/go-build673170913/b001/exe/a.out to cilium-operator: write cilium-operator: no space left on device
Try to remediate that by cleaning up runner disk space before building
Cilium and the HTML docs.
AIL: 0
Signed-off-by: Tobias Klauser <tobias@cilium.io>
The blamed commit intended to add a toleration for the warning occurring in case of failure updating a frontend entry, due to temporary memory pressure. However, it added it to the list of error log exceptions, while the log is actually emitted at warn level. Let's get that fixed. Fixes: 766dbfc ("cli: tolerate occurrences of LB maps cannot allocate memory errors") Signed-off-by: Marco Iorio <marco.iorio@isovalent.com>
Similarly to 9f19618 ("cli: tolerate kvstore nodes GC warning"), let's exclude two more log warnings that can be also triggered due to a race condition in the same logic, and similarly fixed as part of dfb5a38 ("operator: extract and rework kvstore nodes GC logic"). Signed-off-by: Marco Iorio <marco.iorio@isovalent.com>
…aration Remove the legacy Go-side declaration of interface_ifindex. The declaration in bpf/config/global.h introduced in commit 6710178 ("bpf: single source of truth for runtime-subbed node configs") is the source of truth for interface_ifindex where dpgen extracts the runtime config variable. AIL: 0 Signed-off-by: Tobias Klauser <tobias@cilium.io>
The utility function is unused since commit 7f8096e ("bpf, datapath: move IPv6 SNAT exclusion to runtime config"). AIL: 0 Suggested-by: Timo Beckers <timo@isovalent.com> Signed-off-by: Tobias Klauser <tobias@cilium.io>
This helper is simple enough to be inlined. AIL: 0 Signed-off-by: Tobias Klauser <tobias@cilium.io>
We introduced an `extensions` flow field, which was explicitly designed for vendors to enhance or change the behavior of Hubble itself, or Hubble compatible APIs without changing the public API. However, if we try to interact such an extended, Hubble compatible API using the Hubble CLI, JSON marshaling fails with the following error, because it's unable to resolve the type of the extension. ``` json: error calling MarshalJSON for type *flow.Flow: proto: google.protobuf.Any: unable to resolve "example.com/unknown-extension": not found ``` We avoid this error by removing any extension that contains unknown types in the `extensions` field. This makes sure that we can still marshal the flow and just drop unexpected flow extensions. Signed-off-by: Fabian Fischer <fabian.fischer@isovalent.com>
Introduce NodeReconciler and shared identifiers for node-table reconcilers. Use the type for registration APIs so callers do not pass arbitrary strings, and migrate WireGuard to the shared identifier. AIL:3 Related: cilium#41744 Signed-off-by: Jussi Maki <jussi@isovalent.com>
Allow callers to select which registered reconcilers should be refreshed. Mark and wait for only those statuses so unrelated pending work does not delay a targeted refresh. Retain the existing all-reconciler behavior when no names are supplied and reject unregistered reconcilers. AIL:3 Related: cilium#41744 Signed-off-by: Jussi Maki <jussi@isovalent.com>
Move the node checkpoint implementation from manager.go into a dedicated file without changing behavior. This prepares checkpoint ownership to move to the Linux datapath in a separately reviewable patch. AIL:3 Related: cilium#41744 Signed-off-by: Jussi Maki <jussi@isovalent.com>
Move nodes.json restoration, writing, and stale-node pruning out of NodeManager and into the Linux datapath. Watch the node table and wait for all producers to initialize before pruning restored state. Keep the legacy node handler subscription unchanged so the subsequent reconciler conversion can be reviewed independently. AIL:3 Related: cilium#41744 Signed-off-by: Jussi Maki <jussi@isovalent.com>
Simplify checkpoint snapshots with maps and slices helpers and prune live nodes directly from the restored set. Document why failed cleanup entries remain checkpointed for retries and future restarts. AIL:3 Related: cilium#41744 Signed-off-by: Jussi Maki <jussi@isovalent.com>
Convert the Linux node handler from NodeManager callbacks to a reconciler on Table[*Node]. Keep the existing datapath operations behind a small adapter. Declare the Linux reconciler during Hive construction so callers can wait for it immediately. Start its worker after the datapath configuration is available, and retain cluster-size-dependent periodic refreshes. Make node ID allocation safe to retry. AIL:3 Related: cilium#41744 Signed-off-by: Jussi Maki <jussi@isovalent.com>
Use node.Writer.Refresh for IPsec key rotation and subnet topology changes. Remove the Linux datapath implementation of the legacy node.Handler interface now that all node events and refreshes flow through StateDB. AIL:3 Related: cilium#41744 Signed-off-by: Jussi Maki <jussi@isovalent.com>
Expose the per-node encapsulation decision through the 'NodePolicy' object to reduce the exported API surface. Component wanting to override can now depend on 'NodePolicy' instead of 'node.Handler'. AIL:3 Related: cilium#41744 Signed-off-by: Jussi Maki <jussi@isovalent.com>
The NS-traffic conn-disrupt setup picks the node hosting the test-conn-disrupt-server-ns-traffic pod and derives one client deployment per address of that node. It took the first pod of the list and looked its node up in ct.nodes without checking either step, so a missing node gave back a nil *Node and the next dereference of Status.Addresses took the process down with a SIGSEGV. That is what the scheduled ci-ipsec run on EKS hit. The server deployment sat at 0 of 1 available replicas for the whole five-minute wait, and since WaitForDeployment's failure is only recorded through Failf the setup kept going with a pod that was still unscheduled, so its empty NodeName missed in the node map. The panic then replaced the real diagnosis with a stack trace. Guard the three steps that can come up empty: an empty pod list, a pod not yet scheduled, and a node name absent from the node map each return a descriptive error now, which the caller propagates so the setup reports why it could not build the client deployments. https://github.com/cilium/cilium/actions/runs/33718018507/job/100531312300 This commit was prepared with AIL:3. I personally checked the panic stack trace in the linked CI run. Fixes: ba1376a ("cilium-cli: extend no-interrupted-connections to test NodePort from outside") Signed-off-by: André Martins <andre@cilium.io>
setup-gke-cluster reports the zone it used, but conformance-gke keeps deriving the follow-up calls from the zone it asked for. Those are the same today, and stop being the same the moment a fallback zone is configured, at which point the ESP rule and subnet lookups fail and teardown leaks the cluster. AIL:3 Signed-off-by: André Martins <andre@cilium.io>
Capacity starvation does not always fail fast: 13 of the 18659 cluster creations of the last 45 days sat in the GCE instance group manager for its full 35 minute timeout, and net-perf legs cycled zones for 95 to 119 minutes until the job timeout killed them. Bound an attempt to 12 minutes and all attempts to 30. The median creation takes 4.5 minutes and the 99th percentile 7.3, so almost nothing that would have succeeded is aborted. gcloud cannot cancel the operation it submitted and GKE refuses a delete while it runs, so an abandoned attempt is left to the teardown step. AIL:3 Signed-off-by: André Martins <andre@cilium.io>
64 of the 18659 cluster creations of the last 45 days failed on a GCE stockout, all of them in conformance and none in net-perf, every log ending in "All requested zones (us-east4-b ) are out of capacity": conformance-gke never passed fallback-zones, so the loop had a single candidate. Hand it the other zones of the same region, discovered rather than listed. Same region only, because cluster names encode just the config index, so legs of different k8s versions share a name and must not meet in one zone, and each version is pinned to its own region. Skip a zone that does not offer the pinned version, since rollouts are staged per zone and the matrix only checks the zone it asks for. Shuffle the fallbacks, since legs sharing a zone stock out together. AIL:3 Signed-off-by: André Martins <andre@cilium.io>
An attempt abandoned for taking too long keeps provisioning, so the cluster can be in a zone the job has already moved on from, where nothing deletes it. Loop over the candidates instead. The name is unique to the leg, so a delete by name cannot touch another leg's cluster. Also retry a refused delete, which a 503 storm did to 12 clusters on 2026-08-20, and skip the zones where the cluster does not exist, which accounts for 89 of the 173 teardown failures in the window. AIL:3 Signed-off-by: André Martins <andre@cilium.io>
Signed-off-by: Pragalva Sapkota <sapkotapragalva@gmail.com>
Signed-off-by: cilium-renovate[bot] <134692979+cilium-renovate[bot]@users.noreply.github.com>
makeTestEndpointParams left PolicyFetcher unset, so every endpoint it builds carries a nil policy computer. A host-endpoint label update reaches regeneration: ModifyIdentityLabels moves the endpoint into waiting-for-identity, the resolve-identity controller assigns an identity and takes it to ready, and the regeneration that follows dereferences the nil fetcher in waitForPolicyComputationResult. That controller carries a jitter of up to identity-max-jitter, 30 seconds by default, so it usually never fires before teardown; on a slow runner it does, and the package dies with a SIGSEGV that go test blames on whichever test is running, as in https://github.com/cilium/cilium/actions/runs/33652230618. This commit was prepared with AIL:3. I personally checked the diff. Fixes: c597c56 ("endpoint: Plumb policy computer") Signed-off-by: André Martins <andre@cilium.io>
Topology aware routing is covered by unit tests and by the
topology-aware{,-terminating}.txtar script tests, but no kind target
brings up a cluster that exercises it: kind.sh never labels nodes with
zones and no values file enables the feature.
Add `make kind-topology` and `make kind-topology-install-cilium`,
mirroring the kind-egressgw target pair. They create a three worker
cluster and install the locally built images with
loadBalancer.serviceTopology and kube-proxy replacement enabled, so that
Cilium rather than kube-proxy picks the backend. The install target then
applies contrib/testing/topology/: per zone echo deployments, a client
pinned to zone-a and a Service with trafficDistribution: PreferSameZone.
zone-a spans two of the workers so that a backend can sit in the client's
zone but on a different node. With one node per zone the zone-local
backends are exactly the node-local ones, and the environment could not
tell zone based selection apart from node local selection.
Signed-off-by: Gyutae Bae <gyutae.bae@navercorp.com>
Co-authored-by: Minsu Jeon <minsu.jeon@navercorp.com>
Co-authored-by: Siwan Kim <siwan.kim@navercorp.com>
Co-authored-by: Jonghyeon Kim <jong-hyeon.kim@navercorp.com>
Author
|
Closing as requested; this was a manual upstream sync PR. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
mainto match cilium/ciliummain(810 commits; no unique Roblox commits onmain).mainis blocked by branch protection, so this PR is the sync path.Test plan
mainMade with Cursor