Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 9 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,9 +38,11 @@ also sees HTTP paths and headers and lets policy match on method and path
- **One `cgroup_skb/egress` program** at the node's root cgroup v2 sees egress
from **every pod on the node** — DNS, TLS ClientHello SNI, plaintext HTTP, and
new TCP connections — no per-pod sidecar.
- **An `SSL_write` uprobe** reads HTTPS request plaintext **before** encryption,
recovering paths the packet layer can't see. Auto-discovered per container's
libssl, live (no sampling).
- **TLS uprobes** read HTTPS request plaintext **before** encryption, recovering
paths and headers the packet layer can't see. Two probes cover the common cases:
OpenSSL's `SSL_write` (auto-discovered per container's libssl) and Go's
statically-linked `crypto/tls.(*Conn).Write` (auto-discovered per Go binary).
Live, no sampling.
- **Attributed per pod** in-kernel via the originating cgroup id, enriched to
`namespace/name` by a node-scoped Pods informer. The same maps carry policy
verdicts back to the kernel for enforcement.
Expand Down Expand Up @@ -110,9 +112,10 @@ DaemonSet on any node (any CNI, or no Kubernetes at all), watch egress in `log`
mode, and pull it back out with zero effect on connectivity — it never touches the
dataplane.

The headline difference: ebfw reads **HTTPS request paths and headers from an
`SSL_write` uprobe — before encryption, with no proxy and no TLS MITM.** Cilium
needs a terminating Envoy and injected certs to see the same thing.
The headline difference: ebfw reads **HTTPS request paths and headers from TLS
uprobes (OpenSSL `SSL_write` + Go `crypto/tls`) — before encryption, with no proxy
and no TLS MITM.** Cilium needs a terminating Envoy and injected certs to see the
same thing.

ebfw is the lightweight egress firewall + per-pod L7 *visibility* layer; Cilium is
the full networking platform (and does L7 *enforcement* today, which ebfw doesn't
Expand Down
3 changes: 2 additions & 1 deletion ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,5 +34,6 @@ lockdown, log mode, and deferred-dimension negative checks — see
- Pinned maps + an external map-programming controller (the per-node agent programs
its own maps for now).
- TLS/HTTP multi-segment reassembly, TLS 1.3 ECH, cgroup-v1 fallback,
request-body inspection, uprobe coverage beyond OpenSSL-dynamic, multi-kernel CI,
request-body inspection, uprobe coverage beyond OpenSSL-dynamic + Go crypto/tls
(Java, rustls, OpenSSL-static, stripped Go), multi-kernel CI,
HA, and scale tests.
97 changes: 80 additions & 17 deletions bpf/sslsnoop.bpf.c
Original file line number Diff line number Diff line change
@@ -1,17 +1,27 @@
// SPDX-License-Identifier: GPL-2.0
//
// sslsnoop (uprobe PoC).
// sslsnoop (uprobe).
//
// Attaches a uprobe to OpenSSL's SSL_write to capture the *plaintext* of TLS
// traffic before it is encrypted. This is the only non-MITM way to see HTTPS
// request paths (which are encrypted on the wire and therefore invisible to the
// packet-level egress monitor).
// Captures the *plaintext* of TLS traffic before it is encrypted — the only
// non-MITM way to see HTTPS request paths (encrypted on the wire, so invisible
// to the packet-level egress monitor). Two entry points feed one ring buffer:
//
// int SSL_write(SSL *ssl, const void *buf, int num);
// arg1 = SSL* arg2 = buf (plaintext) arg3 = num (length)
// 1. OpenSSL's dynamically-linked SSL_write (curl, nginx, most C/Python/…):
// int SSL_write(SSL *ssl, const void *buf, int num);
// arg1 = SSL* arg2 = buf (plaintext) arg3 = num (length)
//
// The probe copies up to MAX_DATA bytes of the buffer to userspace via a ring
// buffer; HTTP parsing happens in Go.
// 2. Go's statically-linked crypto/tls.(*Conn).Write (any net/http client):
// func (c *Conn) Write(b []byte) (int, error)
// Go does NOT call SSL_write, so OpenSSL's uprobe never sees it. We attach
// directly to the Go symbol instead and read Go's register-based ABIInternal
// (Go >= 1.17): integer/pointer args go in a fixed register sequence, not the
// C ABI, so we cannot use BPF_UPROBE's PT_REGS_PARM* — we read the registers
// by hand per arch. For (*Conn).Write the sequence is (c, b.ptr, b.len, b.cap):
// amd64: RAX, RBX, RCX, RDI arm64: R0, R1, R2, R3
// The *Conn pointer stands in for SSL* as the per-connection key.
//
// Both paths copy up to MAX_DATA bytes of the buffer to userspace via the ring
// buffer; HTTP parsing (HTTP/1.x and HTTP/2 + HPACK) happens in Go.

#include <linux/bpf.h>
#include <bpf/bpf_helpers.h>
Expand Down Expand Up @@ -47,30 +57,30 @@ struct {
__uint(max_entries, 1 << 20); // 1 MiB
} ssl_events SEC(".maps");

// BPF_UPROBE (libbpf >= 1.2, provided by the trixie build image) extracts the
// function arguments from pt_regs.
SEC("uprobe/SSL_write")
int BPF_UPROBE(ssl_write, void *ssl, const void *buf, int num)
// submit_plaintext reserves an event, tags it with the calling task's identity,
// copies up to MAX_DATA bytes of buf, and submits. `key` is the per-connection
// identifier (SSL* for OpenSSL, *Conn for Go). Shared by both uprobes.
static __always_inline int submit_plaintext(__u64 key, const void *buf, __u64 num)
{
if (num <= 0 || buf == 0)
if (num == 0 || buf == 0)
return 0;

struct ssl_event *e = bpf_ringbuf_reserve(&ssl_events, sizeof(*e), 0);
if (!e)
return 0;

e->pid = bpf_get_current_pid_tgid() >> 32;
e->ssl = (__u64)(unsigned long)ssl;
e->ssl = key;
// cgroup v2 id of the calling task, captured here (in process context) so
// attribution doesn't race the process exiting — short-lived TLS clients
// (curl, etc.) are often gone before userspace could read /proc.
e->cgroup_id = bpf_get_current_cgroup_id();
bpf_get_current_comm(&e->comm, sizeof(e->comm));

// Same verifier-friendly clamp as the egress program: __u64 + clamp +
// Verifier-friendly clamp as in the egress program: __u64 + clamp +
// barrier + non-zero guard, so the bound lands on the register passed as
// the read length.
__u64 n = (__u32)num;
__u64 n = num;
if (n > MAX_DATA)
n = MAX_DATA;
barrier_var(n);
Expand All @@ -84,3 +94,56 @@ int BPF_UPROBE(ssl_write, void *ssl, const void *buf, int num)
bpf_ringbuf_submit(e, 0);
return 0;
}

// BPF_UPROBE (libbpf >= 1.2, provided by the trixie build image) extracts the
// function arguments from pt_regs following the C ABI.
SEC("uprobe/SSL_write")
int BPF_UPROBE(ssl_write, void *ssl, const void *buf, int num)
{
if (num <= 0)
return 0;
return submit_plaintext((__u64)(unsigned long)ssl, buf, (__u32)num);
}

// Go's ABIInternal (Go >= 1.17) passes integer/pointer args in a fixed register
// sequence, NOT the C ABI, so BPF_UPROBE's PT_REGS_PARM* would read the wrong
// registers. Read them by hand per arch. go_arg(ctx, i) returns the i-th
// integer-class argument register in ABIInternal order.
#if defined(__TARGET_ARCH_x86)
static __always_inline __u64 go_arg(struct pt_regs *ctx, int i)
{
// Userspace <asm/ptrace.h> names the x86_64 registers with the `r` prefix
// (rax/rbx/…); the bare ax/bx/… names exist only in the kernel/vmlinux
// struct pt_regs, which we don't include. This matches BPF_UPROBE's own
// PT_REGS_PARM* under the UAPI header.
switch (i) {
case 0: return ctx->rax; // RAX
case 1: return ctx->rbx; // RBX
case 2: return ctx->rcx; // RCX
case 3: return ctx->rdi; // RDI
}
return 0;
}
#elif defined(__TARGET_ARCH_arm64)
static __always_inline __u64 go_arg(struct pt_regs *ctx, int i)
{
// arm64 UAPI exposes the register file as struct user_pt_regs (regs[0..30]);
// ABIInternal integer args start at R0.
return ((const struct user_pt_regs *)ctx)->regs[i & 31];
}
#else
static __always_inline __u64 go_arg(struct pt_regs *ctx, int i) { return 0; }
#endif

// crypto/tls.(*Conn).Write(b []byte): args in ABIInternal order are
// (c=*Conn, b.ptr, b.len, b.cap). We key on c and copy b.ptr[0:b.len].
SEC("uprobe/go_tls_write")
int go_tls_write(struct pt_regs *ctx)
{
__u64 conn = go_arg(ctx, 0);
const void *buf = (const void *)go_arg(ctx, 1);
__s64 len = (__s64)go_arg(ctx, 2);
if (len <= 0)
return 0;
return submit_plaintext(conn, buf, (__u64)len);
}
14 changes: 8 additions & 6 deletions docs/comparison.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,10 +34,12 @@ egress-firewall slice of Cilium" (`toFQDNs`, L7 HTTP policy, Hubble flow visibil
1. **No proxy in the data path for L7 visibility.** This is the big one. To see HTTP
method / path / headers on **HTTPS**, Cilium has to terminate TLS through Envoy —
a man-in-the-middle with injected certs sitting in every flow. ebfw recovers the
same plaintext request line by hooking `SSL_write` *before* encryption: no latency
tax, no cert management, no proxy to operate.
*Caveat:* the uprobe is OpenSSL-dynamic only — it misses statically-linked TLS
(Go `crypto/tls`, Java, rustls). Envoy termination catches all of those.
same plaintext request line by hooking the TLS write path *before* encryption: no
latency tax, no cert management, no proxy to operate. Two uprobes cover the common
cases — OpenSSL's dynamic `SSL_write` (curl, nginx, most C/Python/…) and Go's
statically-linked `crypto/tls.(*Conn).Write` (any `net/http` client).
*Caveat:* still misses other statically-linked TLS (Java, rustls, OpenSSL-static,
stripped Go binaries). Envoy termination catches all of those.

2. **Additive, not a commitment.** Adopting Cilium means adopting (or CNI-chaining
into) a network dataplane — a cluster-wide, hard-to-reverse decision. ebfw is a
Expand All @@ -64,8 +66,8 @@ Stated plainly, because the honest comparison matters:
gRPC) today. ebfw evaluates those dimensions for log + metrics but **does not yet
drop on them** — that needs the terminating proxy + TLS MITM we've deliberately
avoided. It's on the [roadmap](https://github.com/dvrkn/ebfw/blob/main/ROADMAP.md).
- **Universal TLS coverage.** Envoy termination sees every TLS library; our uprobe
sees OpenSSL-dynamic only.
- **Universal TLS coverage.** Envoy termination sees every TLS library; our uprobes
cover OpenSSL-dynamic + Go `crypto/tls` (not Java, rustls, or stripped/static).
- **Breadth & maturity.** Ingress policy, encryption, load balancing, multi-cluster,
identity-based scaling, HA, multi-kernel CI, scale tests, CNCF-graduated. ebfw is
early: IPv6 enforcement is incomplete, there's no map pinning yet, and it's
Expand Down
6 changes: 4 additions & 2 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -151,8 +151,10 @@ in the event lines, not in metric labels.
and IPv6, but an IPv6 packet whose next-header is not TCP/UDP directly (a
hop-by-hop, routing, fragment, or destination-options header) is skipped rather
than walked. These are rare on normal egress.
- **HTTPS paths need OpenSSL-dynamic.** Statically-linked TLS has no `libssl.so`
to hook — notably Go (`crypto/tls`, often stripped), Java, rustls.
- **HTTPS paths need OpenSSL-dynamic or Go crypto/tls.** Two uprobes cover these:
OpenSSL's `SSL_write` (via `libssl.so`) and Go's `crypto/tls.(*Conn).Write` (via
the Go binary's symbol table). Still uncovered: Java, rustls, statically-linked
OpenSSL, and *stripped* Go binaries (`-ldflags "-s -w"` removes the symbol).
- **Single segment** — DNS/TLS/HTTP are parsed from the first packet/segment only;
larger ClientHellos/requests are truncated. Reassembly is future work.
- **Request bodies are a stub** (`EBFW_INSPECT_BODY` does nothing yet).
Expand Down
5 changes: 4 additions & 1 deletion docs/tests.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,10 @@ Linux host, eBPF, root. Covers visibility — DNS, TLS SNI, HTTP/1.x + HTTP/2
request paths, headers — plus internal-traffic filtering, IPv4 + IPv6 packet
decoding, JSON output, and enforcement: a CIDR deny drops the connection, a domain
deny blocks via DNS→IP learning, and `log` mode annotates the verdict without
dropping.
dropping. It also builds a native Go `net/http` client (statically-linked
`crypto/tls`, invisible to the OpenSSL uprobe) and asserts ebfw recovers its host,
path, and a custom header via the `crypto/tls.(*Conn).Write` uprobe — self-skipping
where no Go toolchain is present.

## Kubernetes e2e — `test/crd.sh`

Expand Down
7 changes: 7 additions & 0 deletions internal/metrics/metrics.go
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,13 @@ var (
Help: "Number of libssl SSL_write uprobes currently attached.",
})

// GoUprobesAttached is the number of Go crypto/tls (*Conn).Write uprobes
// attached — one per unique Go binary that links crypto/tls statically.
GoUprobesAttached = promauto.NewGauge(prometheus.GaugeOpts{
Name: "ebfw_go_uprobe_attached",
Help: "Number of Go crypto/tls (*Conn).Write uprobes currently attached.",
})

// EnforcementDecisionsTotal counts policy verdicts applied to connection-level
// events, by action (allow/deny/modify) and mode (log/enforce).
EnforcementDecisionsTotal = promauto.NewCounterVec(prometheus.CounterOpts{
Expand Down
136 changes: 135 additions & 1 deletion internal/sslsnoop/sslsnoop.go
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@ package sslsnoop
import (
"bytes"
"context"
"debug/elf"
"encoding/binary"
"errors"
"fmt"
Expand Down Expand Up @@ -72,7 +73,8 @@ func Run(ctx context.Context, cfg *config.Config, filter *config.Filter, resolve
}()

go discoverLoop(ctx, objs.SslWrite)
log.Printf("ebfw sslsnoop: auto-discovering libssl across the node (live capture)")
go discoverGoLoop(ctx, objs.GoTlsWrite)
log.Printf("ebfw sslsnoop: auto-discovering libssl + Go crypto/tls across the node (live capture)")

t := newTracker(filter, cfg.Inspect, resolver, sink)
for {
Expand Down Expand Up @@ -229,6 +231,138 @@ func discoverLoop(ctx context.Context, prog *ebpf.Program) {
}
}

// goTLSWriteSym is the Go symbol we attach to for statically-linked TLS. It is
// present in the .symtab of any non-stripped Go binary that imports crypto/tls
// (the default `go build` keeps the symbol table).
const goTLSWriteSym = "crypto/tls.(*Conn).Write"

// discoverGoLoop attaches the go_tls_write uprobe to the crypto/tls.(*Conn).Write
// symbol of every unique Go binary on the node, rescanning on the same cadence
// as the libssl loop to pick up new containers. Symbol presence is immutable per
// inode, so each binary's ELF is parsed at most once.
func discoverGoLoop(ctx context.Context, prog *ebpf.Program) {
attached := map[string]link.Link{}
hasSym := map[string]bool{} // inode -> ELF parsed, symbol present?
failed := map[string]bool{} // inode -> attach error already logged
defer func() {
for _, l := range attached {
l.Close()
}
}()
tick := time.NewTicker(discoverInterval)
defer tick.Stop()
for {
for key, path := range discoverExecutables() {
if _, ok := attached[key]; ok {
continue
}
present, ok := hasSym[key]
if !ok {
// First sighting of this binary: parse its symbol table. A parse
// error (e.g. the pid exited, path now stale) is left uncached so
// the next tick retries via a different live pid.
p, err := hasGoTLSSymbol(path)
if err != nil {
continue
}
hasSym[key] = p
present = p
}
if !present {
continue // not a Go crypto/tls binary
}
ex, err := link.OpenExecutable(path)
if err != nil {
if !failed[key] {
log.Printf("ebfw sslsnoop: open %s: %v", path, err)
failed[key] = true
}
continue
}
up, err := ex.Uprobe(goTLSWriteSym, prog, nil)
if err != nil {
if !failed[key] {
log.Printf("ebfw sslsnoop: go uprobe %s: %v", path, err)
failed[key] = true
}
continue
}
delete(failed, key)
attached[key] = up
metrics.GoUprobesAttached.Inc()
log.Printf("ebfw sslsnoop: attached %s uprobe via %s [inode %s]", goTLSWriteSym, path, key)
}
select {
case <-ctx.Done():
return
case <-tick.C:
}
}
}

// discoverExecutables returns "dev:inode" -> /proc/<pid>/exe for every unique
// process image on the node, deduped by inode (each physical binary attached
// once). Unlike a shared library — one file per mount namespace — every process
// has its own executable, so we must stat every pid's exe rather than dedup by
// mount namespace. The /proc/<pid>/exe magic symlink is openable for ELF parsing
// and uprobe attachment across namespaces.
func discoverExecutables() map[string]string {
out := map[string]string{}
// Exclude the agent's own binary: ebfw links crypto/tls (via client-go), so
// it carries goTLSWriteSym; attaching to ourselves would just capture the
// agent's own API-server traffic.
var self syscall.Stat_t
haveSelf := syscall.Stat("/proc/self/exe", &self) == nil

procs, err := os.ReadDir("/proc")
if err != nil {
return out
}
for _, p := range procs {
pid, err := strconv.Atoi(p.Name())
if err != nil {
continue
}
exe := fmt.Sprintf("/proc/%d/exe", pid)
var st syscall.Stat_t
if err := syscall.Stat(exe, &st); err != nil {
continue
}
if haveSelf && st.Dev == self.Dev && st.Ino == self.Ino {
continue
}
key := fmt.Sprintf("%d:%d", st.Dev, st.Ino)
if _, ok := out[key]; !ok {
out[key] = exe
}
}
return out
}

// hasGoTLSSymbol reports whether the ELF at path defines goTLSWriteSym in its
// symbol table (i.e. it is a non-stripped Go binary linking crypto/tls). A
// missing symbol table is not an error — it just means "not a match".
func hasGoTLSSymbol(path string) (bool, error) {
f, err := elf.Open(path)
if err != nil {
return false, err
}
defer f.Close()
syms, err := f.Symbols()
if err != nil {
if errors.Is(err, elf.ErrNoSymbols) {
return false, nil
}
return false, err
}
for i := range syms {
if syms[i].Name == goTLSWriteSym {
return true, nil
}
}
return false, nil
}

// discoverLibssl finds the libssl shared object in every container's root
// filesystem via /proc/<pid>/root and returns "dev:inode" -> path. Dedup by
// mount namespace keeps it to one scan per container; dedup by inode means each
Expand Down
Loading