diff --git a/README.md b/README.md index c114ca0..c9eedff 100644 --- a/README.md +++ b/README.md @@ -12,9 +12,33 @@ func main() { } ``` -Guest startup is internal to the library. A new guest executes Go runtime and package initialization, then enters the imported closure without executing the application's `main`. Initialization side effects must be appropriate there. Captured objects must be exclusively owned until `Run` returns. The function must finish its own goroutines before returning. Concurrent and nested host calls are supported when their captured graphs are independent. +Guest startup is internal to the library. A new guest executes Go runtime and package initialization, then enters the imported closure without executing the application's `main`. Initialization side effects must be appropriate there. Captured objects must be exclusively owned until `Run` returns. Goroutines must stop accessing captured objects before the callback returns; a persistent Process may retain goroutines that only use guest-owned state. Concurrent and nested host calls are supported when their captured graphs are independent. -`sandbox.Run` uses one global default Sandbox, initialized on first use and retained for the host process lifetime. Each explicit `Sandbox` likewise starts its Kernel on its first `Run` and reuses it until `Close`. Every call creates a fresh guest process, PID namespace, filesystem namespace, FD table and snapshot, and transfers the complete type and object graph. Guest processes are not pooled. `Close` rejects new calls, terminates active guests and waits for their calls to finish; repeated `Close` calls are safe. Do not call it from that Sandbox's inspector. Do not copy a used Sandbox or change its fields while calls are active; `Library` must remain unchanged after first use. +`sandbox.Run` uses one global default Sandbox, initialized on first use and retained for the host process lifetime. Each explicit `Sandbox` likewise reuses its Kernel until `Close`. These one-shot entry points create and close a fresh guest for each call. `Sandbox.Close` rejects new calls, terminates its guests and waits for active calls to finish. Do not call Close from the corresponding inspector, copy a used Sandbox, or change its configuration while calls are active; `Library` remains fixed after first use. + +## Persistent processes + +Use one Process when callbacks need the same native package globals, open guest files or filesystem state: + +```go +p := sandbox.NewProcess(sandbox.ProcessOptions{ + Env: os.Environ(), +}) +defer p.Close() + +if err := p.Run(func() { f.OnRequire(proj, deps) }); err != nil { + return err +} +if err := p.Run(func() { f.OnBuild(ctx) }); err != nil { + return err +} +``` + +`NewProcess` accepts zero or one `ProcessOptions` value. It uses the default shared Kernel and starts its guest on the first Run. Options contain `Mounts`, `Env` and `Inspect`; nil Env inherits the host environment at startup. The guest's environment and mounts then persist. Each successful Run writes captured changes back, pauses this process's guest tasks and leaves native global state in the guest. Calls on one Process execute serially; independent processes may run concurrently. Existing object IDs are retained across calls, including when a guest global keeps a pointer to a captured object. Close terminates the process and any active call, waits for cleanup, and is idempotent. Failed execution or a failed transfer closes the process; it is never silently restarted. + +After ten seconds idle, Sentry flushes eligible private pages and asks the host to reclaim their resident storage. The process, virtual mappings and kernel resources stay alive; subsequent accesses fault pages back at the same addresses. The application MemoryFile is backed by an unlinked temporary file. Place the host temporary directory on a disk filesystem for reclamation; a tmpfs-backed file cannot provide this disk-eviction behavior. Eviction failures are logged and do not discard guest data. Shared/COW mappings and Sentry's own management memory are retained. This is live-process page eviction, not a checkpoint that survives host shutdown. + +The eviction adapter reads the pinned gVisor version's private mapping metadata under its existing lock. It does not modify gVisor or parse diagnostic text. Each Run still uses a fresh memfd, and results remain write-sealed before host decoding. The one-shot `sandbox.Run(fn)` creates a Process, runs it and closes it automatically. ## Modules @@ -81,8 +105,8 @@ err := s.Run(func() { f.OnBuild(ctx) }) | Type | Source and lifetime | | --- | --- | | `bind` | Host directory, served through DirectFS. Writes persist in that host directory. Host permissions still apply. | -| `tmpfs` | New Sentry filesystem for each call. Supports options such as `size`, `mode`, `uid` and `gid`; contents disappear when the call ends. | -| `proc` | Process information from this call's PID namespace. | +| `tmpfs` | New Sentry filesystem for each process. Supports options such as `size`, `mode`, `uid` and `gid`; contents persist until the process closes. | +| `proc` | Process information from the guest's PID namespace. | | `overlay` | Combines already visible guest paths using `lowerdir` and optional `upperdir`. Writes go to the upper layer, whose filesystem determines persistence. | An empty `Mounts` retains the default read-only host `/` and guest `/proc`. A nonempty list replaces all defaults. Its first entry must mount `bind` or `tmpfs` at `/`; subsequent mounts are applied in order, so parents and overlay layers must precede their users. Duplicate targets and unsupported types return an error. Missing directory mountpoints are prepared through Sentry's synthetic-mountpoint support, which still requires a writable parent mount. For a read-only parent, prepare the target directory beforehand or place new mountpoints under a writable tmpfs. Bind sources currently must be directories. diff --git a/bridge.h b/bridge.h index 3ed2877..8db4f01 100644 --- a/bridge.h +++ b/bridge.h @@ -21,7 +21,8 @@ struct syscall_event { typedef void (*inspect_fn)(uintptr_t, struct syscall_event *); typedef int (*create_sentry_fn)(uintptr_t *, char *, size_t); typedef int (*close_sentry_fn)(uintptr_t, char *, size_t); -// Each run receives a Kernel handle and startup JSON (guest, mounts and env). +// Dispatches create/run/close for a process ID within a Kernel. +// Creation supplies guest/mount/env configuration; each run supplies a fresh image fd. typedef int (*run_sentry_fn)(uintptr_t, char *, int, uintptr_t, uintptr_t, uintptr_t, inspect_fn, char *, size_t); diff --git a/guest_linux.go b/guest_linux.go index de1ae69..4f87b60 100644 --- a/guest_linux.go +++ b/guest_linux.go @@ -6,6 +6,7 @@ import ( "context" "errors" "fmt" + "io" "os" "runtime" @@ -30,23 +31,39 @@ func runGuest() (err error) { err = fmt.Errorf("sandbox guest panicked: %v", value) } }() + control := os.NewFile(4, "sandbox-control") + defer control.Close() defer unix.Close(3) - data, unmap, err := readStateImage(3, 0) - if err != nil { - return err - } - defer func() { err = errors.Join(err, unmap()) }() var graph state.State var fn func() - ctx := context.Background() - if _, err := graph.Load(ctx, data, &fn); err != nil { - return err - } - runtime.GC() - fn() - _, err = writeStateImage(3, int64(len(data))+8, &graph, &fn) - if err != nil { - return err + var command [1]byte + for { + if _, err := io.ReadFull(control, command[:]); err != nil { + if errors.Is(err, io.EOF) { + return nil + } + return err + } + if command[0] != 1 { + return errors.New("sandbox: invalid process command") + } + data, unmap, err := readStateImage(3, 0) + if err != nil { + return err + } + resultOffset := int64(len(data)) + 8 + _, loadErr := graph.Load(context.Background(), data, &fn) + err = errors.Join(loadErr, unmap()) + if err != nil { + return err + } + runtime.GC() + fn() + if _, err := writeStateImage(3, resultOffset, &graph, &fn); err != nil { + return err + } + if _, err := control.Write(command[:]); err != nil { + return err + } } - return nil } diff --git a/host_linux.go b/host_linux.go index d96e2ed..6fd2f87 100644 --- a/host_linux.go +++ b/host_linux.go @@ -80,82 +80,108 @@ func sandboxInspect(owner C.uintptr_t, event *C.struct_syscall_event) { } } -// Run executes fn in a fresh guest running the same ELF. The guest enters the -// closure automatically after package initialization, without running main. -// Capture mutations are committed only after a successful guest exit. -// Captures must be exclusively owned for the duration of Run. Calls may overlap -// when their captured graphs are independent. Configuration must not change -// until all active calls return. +type processState struct { + kernel uintptr + inspection cgo.Handle + observer *inspection + graph state.State + fn func() +} + +// Run executes fn in a new guest and closes it after writing captures back. func (s *Sandbox) Run(fn func()) (err error) { + p := NewProcess(ProcessOptions{Mounts: s.Mounts, Env: s.Env, Inspect: s.Inspect}) + p.owner = s + defer func() { err = errors.Join(err, p.Close()) }() + return p.Run(fn) +} + +// Run executes fn in this process and pauses the guest before returning. +// Calls on one Process are serialized. Captures must be exclusively owned +// until Run returns. A failed transfer or guest execution closes the process. +func (p *Process) Run(fn func()) (err error) { if fn == nil { - return fmt.Errorf("sandbox: nil function") + return errors.New("sandbox: nil function") } - handleID, err := s.acquire() - if err != nil { + p.runMu.Lock() + p.mu.Lock() + if p.closed || p.failure != nil || p.owner == nil { + err = p.failure + if err == nil { + err = errors.New("sandbox: process is closed or uninitialized") + } + p.mu.Unlock() + p.runMu.Unlock() return err } - defer s.active.Done() - mainPC, err := guestMain() + p.runs.Add(1) + p.mu.Unlock() + transfer := false + defer func() { + if err != nil && transfer { + p.mu.Lock() + p.failure = err + p.mu.Unlock() + } + p.runs.Done() + p.runMu.Unlock() + if err != nil && transfer { + err = errors.Join(err, p.Close()) + } + }() + kernelID, err := p.owner.acquire() if err != nil { return err } - entryPC := reflect.ValueOf(guestEntry).Pointer() + defer p.owner.active.Done() fd, err := newStateImage() if err != nil { return err } defer unix.Close(fd) - var graph state.State - ctx := context.Background() - resultOffset, err := writeStateImage(fd, 0, &graph, &fn) + p.native.fn = fn + resultOffset, err := writeStateImage(fd, 0, &p.native.graph, &p.native.fn) if err != nil { return fmt.Errorf("sandbox export: %w", err) } + transfer = true runtime.GC() - executable, err := os.Executable() - if err != nil { - return err + p.mu.Lock() + if p.closed { + p.mu.Unlock() + return errors.New("sandbox: process is closed") } - mounts := s.Mounts - if len(mounts) == 0 { - mounts = []Mount{ - {Type: "bind", Source: "/", Target: "/", Options: []string{"ro"}}, - {Type: "proc", Target: "/proc"}, + if p.native.kernel == 0 { + p.native.kernel = kernelID + p.native.observer = &inspection{fn: p.options.Inspect} + if p.options.Inspect != nil { + p.native.inspection = cgo.NewHandle(p.native.observer) + } + err = p.call("create", fd) + if err == nil { + p.owner.mu.Lock() + if p.owner.processes == nil { + p.owner.processes = make(map[uint64]*Process) + } + p.owner.processes[p.id] = p + p.owner.mu.Unlock() } } - env := s.Env - if env == nil { - env = os.Environ() - } - config, err := json.Marshal(struct { - Guest string `json:"guest"` - Mounts []Mount `json:"mounts"` - Env []string `json:"env"` - }{executable, mounts, env}) + p.mu.Unlock() if err != nil { - return fmt.Errorf("sandbox configuration: %w", err) - } - cConfig := C.CString(string(config)) - defer C.free(unsafe.Pointer(cConfig)) - i := &inspection{fn: s.Inspect} - var handle cgo.Handle - if s.Inspect != nil { - handle = cgo.NewHandle(i) - defer handle.Delete() + return err } - var message [4096]C.char - code := C.sandbox_run(C.uintptr_t(handleID), cConfig, C.int(fd), C.uintptr_t(mainPC), C.uintptr_t(entryPC), C.uintptr_t(handle), &message[0], C.size_t(len(message))) - if code != 0 { - return fmt.Errorf("sandbox Sentry: %s", C.GoString(&message[0])) + if err := p.call("run", fd); err != nil { + return err } - i.mu.Lock() - inspectionErr := i.err - i.mu.Unlock() + p.native.observer.mu.Lock() + inspectionErr := p.native.observer.err + p.native.observer.mu.Unlock() if inspectionErr != nil { return inspectionErr } - // Seal before either decode pass, including against any fd the guest - // transferred elsewhere. Existing writable mappings make this fail. + // Each call gets a new memfd. The previous result stays sealed even if the + // resident guest retained a duplicate descriptor across calls. if _, err := unix.FcntlInt(uintptr(fd), unix.F_ADD_SEALS, unix.F_SEAL_WRITE|unix.F_SEAL_SEAL); err != nil { return fmt.Errorf("sandbox result seal: %w", err) } @@ -164,14 +190,82 @@ func (s *Sandbox) Run(fn func()) (err error) { return fmt.Errorf("sandbox result: %w", err) } defer func() { err = errors.Join(err, unmap()) }() - if _, err := graph.Load(ctx, data, &fn); err != nil { + if _, err := p.native.graph.Load(context.Background(), data, &p.native.fn); err != nil { return fmt.Errorf("sandbox import: %w", err) } - graph = state.State{} runtime.GC() return nil } +func (p *Process) call(operation string, fd int) error { + config := struct { + Operation string `json:"operation"` + Process uint64 `json:"process"` + Guest string `json:"guest,omitempty"` + Mounts []Mount `json:"mounts,omitempty"` + Env []string `json:"env,omitempty"` + }{Operation: operation, Process: p.id} + var mainPC, entryPC uintptr + if operation == "create" { + var err error + mainPC, err = guestMain() + if err != nil { + return err + } + entryPC = reflect.ValueOf(guestEntry).Pointer() + config.Guest, err = os.Executable() + if err != nil { + return err + } + config.Mounts = p.options.Mounts + if len(config.Mounts) == 0 { + config.Mounts = []Mount{{Type: "bind", Source: "/", Target: "/", Options: []string{"ro"}}, {Type: "proc", Target: "/proc"}} + } + config.Env = p.options.Env + if config.Env == nil { + config.Env = os.Environ() + } + } + data, err := json.Marshal(config) + if err != nil { + return fmt.Errorf("sandbox configuration: %w", err) + } + cConfig := C.CString(string(data)) + defer C.free(unsafe.Pointer(cConfig)) + var message [4096]C.char + if C.sandbox_run(C.uintptr_t(p.native.kernel), cConfig, C.int(fd), C.uintptr_t(mainPC), C.uintptr_t(entryPC), C.uintptr_t(p.native.inspection), &message[0], C.size_t(len(message))) != 0 { + return fmt.Errorf("sandbox Sentry: %s", C.GoString(&message[0])) + } + return nil +} + +// Close terminates this process, including an active Run, and releases its +// resources. It is idempotent. Do not call it from this process's inspector. +func (p *Process) Close() error { + p.closeOnce.Do(func() { + p.mu.Lock() + p.closed = true + started := p.native.kernel != 0 + p.mu.Unlock() + if started { + p.closeErr = p.call("close", -1) + } + p.runs.Wait() + if p.native.inspection != 0 { + p.native.inspection.Delete() + p.native.inspection = 0 + } + p.native.graph = state.State{} + p.native.fn = nil + if p.owner != nil { + p.owner.mu.Lock() + delete(p.owner.processes, p.id) + p.owner.mu.Unlock() + } + }) + return p.closeErr +} + func (s *Sandbox) acquire() (uintptr, error) { s.mu.Lock() defer s.mu.Unlock() @@ -237,6 +331,15 @@ func (s *Sandbox) Close() error { } } s.active.Wait() + s.mu.Lock() + processes := make([]*Process, 0, len(s.processes)) + for _, p := range s.processes { + processes = append(processes, p) + } + s.mu.Unlock() + for _, p := range processes { + err = errors.Join(err, p.Close()) + } s.closeErr = err close(s.closeDone) return err diff --git a/process.go b/process.go new file mode 100644 index 0000000..eee0cb3 --- /dev/null +++ b/process.go @@ -0,0 +1,59 @@ +package sandbox + +import ( + "errors" + "sync" + "sync/atomic" +) + +// ProcessOptions configures the filesystem, environment and syscall observer +// of a Process. Configuration is fixed when the process first starts. +type ProcessOptions struct { + Mounts []Mount + Env []string + Inspect func(*Syscall) +} + +// Process executes successive callbacks in one guest. Captured objects are +// written back after each successful Run; native package globals remain in the +// guest. Idle guests are paused, and their private pages may be reclaimed after +// ten seconds. Close releases the process and its retained object identities. +// Do not copy a Process. Use NewProcess to create one. +type Process struct { + options ProcessOptions + owner *Sandbox + id uint64 + + runMu sync.Mutex + mu sync.Mutex + runs sync.WaitGroup + closeOnce sync.Once + closed bool + failure error + closeErr error + native processState +} + +var nextProcessID atomic.Uint64 + +// NewProcess creates a process using the shared default Sandbox. It accepts +// zero or one options value. Startup is lazy; the first Run reports startup or +// configuration errors. A nil Env inherits the host environment at startup. +func NewProcess(opts ...ProcessOptions) *Process { + p := &Process{owner: &defaultSandbox, id: nextProcessID.Add(1)} + if len(opts) > 1 { + p.failure = errors.New("sandbox: NewProcess accepts at most one options value") + return p + } + if len(opts) == 1 { + p.options = opts[0] + if opts[0].Env != nil { + p.options.Env = append([]string{}, opts[0].Env...) + } + p.options.Mounts = append([]Mount(nil), opts[0].Mounts...) + for i := range p.options.Mounts { + p.options.Mounts[i].Options = append([]string(nil), opts[0].Mounts[i].Options...) + } + } + return p +} diff --git a/sandbox.go b/sandbox.go index a3df62a..675c919 100644 --- a/sandbox.go +++ b/sandbox.go @@ -2,7 +2,10 @@ // captured object graph back to the caller. See README.md for transfer limits. package sandbox -import "sync" +import ( + "errors" + "sync" +) // Syscall is a trapped guest syscall, observed after Context.Switch returns and // before Sentry dispatches it. Addresses in Args belong to the guest. @@ -44,9 +47,14 @@ type Sandbox struct { closed bool closeDone chan struct{} closeErr error + processes map[uint64]*Process } var defaultSandbox Sandbox -// Run executes fn using a process-lifetime default Sandbox. -func Run(fn func()) error { return defaultSandbox.Run(fn) } +// Run executes fn in a one-shot Process using the shared default Sandbox. +func Run(fn func()) (err error) { + p := NewProcess() + defer func() { err = errors.Join(err, p.Close()) }() + return p.Run(fn) +} diff --git a/sentry/README.md b/sentry/README.md index db5c253..3c69617 100644 --- a/sentry/README.md +++ b/sentry/README.md @@ -37,21 +37,23 @@ int CloseSandbox(uintptr_t kernel, char *message, size_t capacity); - `CreateSandbox` starts one Kernel and returns its opaque handle. `RunSandbox` creates a fresh process under that Kernel, with independent PID and mount namespaces, FD table and inspection state. Calls may overlap. The internal task identity is inherited by guest children so inspection and cleanup remain scoped to the originating Run. - `CloseSandbox` rejects further calls, terminates tasks, waits for active runs and releases the Kernel. The caller must wait for its inspector callbacks to return; closing a Kernel synchronously from one of its own callbacks would deadlock. - `config` is a NUL-terminated JSON object containing `guest` (the executable's absolute guest path), `mounts` and `env` (an array of `KEY=value` strings). Sentry starts the executable with that path as `argv[0]` and no extra arguments. Mount entries contain `type`, optional `source`, `target` and optional `options` (an array of strings). Environment entries are passed directly to the new process before runtime initialization; NUL is rejected. Missing, null or empty `env` gives an empty environment at this C entry. The Go host implements `Sandbox.Env == nil` inheritance by explicitly sending its own `os.Environ()`. -- `image_fd` is a caller-owned descriptor imported as guest fd 3. The library does not interpret its contents. Guest fd 0, 1 and 2 are imported from host stdin, stdout and stderr. +- `image_fd` is a caller-owned descriptor imported as guest fd 3. The library does not interpret its contents. Guest fd 0, 1 and 2 are imported from host stdin, stdout and stderr. Creation also installs an internal control socket at fd 4. The guest reads a byte with value 1 before loading fd 3 and sends the same byte after publishing its result. Fds 3 and 4 are close-on-exec. - `main_pc` is the guest virtual address to redirect; the caller supplies its `main.main` address. The caller must verify that the symbol contains at least 5 bytes on AMD64 or 4 bytes on ARM64. Before starting guest tasks, the library writes a relative branch into this private executable mapping using Sentry's existing memory manager. -- `entry_pc` is the guest virtual address of the caller's private, non-capturing Go `func()` startup entry. It runs after Go package initialization and returns after exporting closure results. The caller owns ELF symbol resolution, guest code and value reconstruction. The library checks branch range and alignment; it does not interpret the closure image. Unsupported branch layouts fail before creating the guest. +- `entry_pc` is the guest virtual address of the caller's private, non-capturing Go `func()` startup entry. It runs after Go package initialization and waits for successive task commands until the process closes. The caller owns ELF symbol resolution, guest code and value reconstruction. The library checks branch range and alignment; it does not interpret the closure image. Unsupported branch layouts fail before creating the guest. - `owner` is an opaque integer passed unchanged to `inspect`. A Go caller can use a `cgo.Handle` owned by its own runtime. - `inspect` is an optional synchronous callback. It receives a borrowed `syscall_event` containing the syscall number, name, six arguments and a memory mapping callback. The caller owns argument parsing and may change the registers directly. A null callback skips inspection setup. Guest pointer arguments must not be dereferenced in the host. -- `message` is a caller-owned writable error buffer of `capacity` bytes. For nonzero capacity, errors are truncated to at most `capacity - 1` bytes and NUL-terminated. No bytes are written at zero capacity. A nonzero result indicates an error; zero means that the guest exited successfully. +- `message` is a caller-owned writable error buffer of `capacity` bytes. For nonzero capacity, errors are truncated to at most `capacity - 1` bytes and NUL-terminated. No bytes are written at zero capacity. A nonzero result indicates an error; zero means that the process operation completed successfully. -All supplied strings, buffers and callback state must remain valid until `RunSandbox` returns. Each event, its name and its context handle are borrowed only for the duration of `inspect`. Load one library per host process and keep it loaded: its Go runtime and Systrap workers retain executable code for the process lifetime, even after all Kernels are closed. +Configuration strings and error buffers must remain valid until the operation returns. The inspector callback and its owner state must remain valid until the created process is closed, including between Run calls. Each event, its name and its context handle are borrowed only for the duration of `inspect`. Load one library per host process and keep it loaded: its Go runtime and Systrap workers retain executable code for the process lifetime, even after all Kernels are closed. The lifecycle entry points and leading Kernel handle change the C ABI. Backends through `sentry/v0.5.0` are incompatible and lack `CreateSandbox`/`CloseSandbox`. Build both modules from the same source revision. Go module version selection does not validate a library loaded with `dlopen`. -Example startup configuration: +Example process creation configuration: ```json { + "operation": "create", + "process": 1, "guest": "/opt/llar/llar", "env": ["PATH=/usr/bin:/bin", "LANG=C"], "mounts": [ @@ -65,7 +67,7 @@ Example startup configuration: The backend requires an explicit list beginning with a `bind` or `tmpfs` root at `/`. The Go host supplies the default read-only `/` and guest `/proc` when its `Mounts` is empty. Bind sources must be absolute host directory paths. Targets are absolute guest paths; mounts are applied in order, and duplicate targets are rejected. Parent mounts and overlay layers must precede their users. Missing directory mountpoints use Sentry's synthetic-mountpoint support and require a writable parent mount; targets under a read-only parent must already exist. -Registered types are `bind` (translated to gofer with DirectFS), `tmpfs`, `proc` and `overlay`. Common options are `ro`/`rw`, `noexec`/`exec`, `nosuid`/`suid` and `noatime`/`atime`, with the last option in a pair taking precedence. Tmpfs and overlay filesystem options are passed to Sentry as mount data; unsupported options fail. For overlay, `lowerdir` and `upperdir` refer to paths in the guest namespace, not host paths. An upper tmpfs makes changes temporary; use writable binds for persistent outputs. Mount namespaces, tmpfs contents and bind connections are released after the call, including partial setup failures. +Registered types are `bind` (translated to gofer with DirectFS), `tmpfs`, `proc` and `overlay`. Common options are `ro`/`rw`, `noexec`/`exec`, `nosuid`/`suid` and `noatime`/`atime`, with the last option in a pair taking precedence. Tmpfs and overlay filesystem options are passed to Sentry as mount data; unsupported options fail. For overlay, `lowerdir` and `upperdir` refer to paths in the guest namespace, not host paths. An upper tmpfs makes changes temporary; use writable binds for persistent outputs. Mount namespaces, tmpfs contents and bind connections live until the process closes; partial setup failures also release them. The event's `mmap(context, address, size, memory)` callback allocates anonymous temporary guest pages using `MMap` and pins them. Address `0` requests `size` zeroed bytes without a source; a nonzero address initializes the pages using `CopyIn`. Sentry chooses a vacant guest address. The returned `syscall_memory` contains a host `data` pointer directly mapping those pages, their guest `address`, and the initialized `length`. The host and guest addresses need not match. Editing `data` immediately changes the temporary pages without changing the original guest bytes. There is no caller-owned staging buffer or write/commit callback; initialization from a source copies bytes and is not COW. Address-zero allocation requires `sentry/v0.3.0` or later; it is not supported by `sentry/v0.2.0`. @@ -89,7 +91,13 @@ guest syscall -> next Context.Switch resumes the guest ``` -The hook is implemented in [platform_linux.go](platform_linux.go). It does not modify gVisor's sysmsg queues, futex handoff or kernel syscall implementations. [run_linux.go](run_linux.go) creates a new Sentry kernel/guest for each call; the Systrap platform is retained between calls. +The hook is implemented in [platform_linux.go](platform_linux.go). It does not modify gVisor's sysmsg queues, futex handoff or kernel syscall implementations. [run_linux.go](run_linux.go) owns persistent guest processes within each Kernel. A completed run stops only its process; other guests remain runnable. The Systrap platform is retained between Kernels. + +## Process operations and idle memory + +`RunSandbox` dispatches `create`, `run` and `close` operations from the configuration JSON, keyed by `process` within its Kernel. Creation installs the guest entry and inspector. Run replaces fd 3 while the guest is paused, wakes it, exchanges a command/completion byte on fd 4 and waits for its tasks to stop again. Close interrupts active control I/O, terminates the process and waits for its threads, mappings and filesystem services to release. CloseSandbox closes all of its processes before releasing the Kernel. This operation protocol requires host and backend from the same source revision, even when an older backend exports the same C symbol names. + +Each process schedules idle eviction ten seconds after completing a run. A new run cancels the previous idle timer; timer callbacks verify their identity under the process lock so an old timer cannot evict a later call. The Kernel's application MemoryFile uses an unlinked temporary backing file. Eviction holds the selected memory manager's mapping lock, flushes existing private non-COW ranges, removes platform mappings and requests file-cache reclamation. It retains VMA/PMA and thread state, and never calls Decommit on live data. The ordinary fault path reloads pages after resume. The pinned gVisor dependency remains unchanged; `process_memory_linux.go` adapts its private mapping metadata using reflection and its existing methods. Host temporary storage must be disk-backed for this reclamation to work. This mechanism retains kernel/FD/VFS state and is not a standalone checkpoint. ## Runtime Conditions diff --git a/sentry/boot_linux.go b/sentry/boot_linux.go index e13e204..82e9ecf 100644 --- a/sentry/boot_linux.go +++ b/sentry/boot_linux.go @@ -26,8 +26,8 @@ import ( "runtime" "sync" + "golang.org/x/sys/unix" "gvisor.dev/gvisor/pkg/cpuid" - "gvisor.dev/gvisor/pkg/memutil" "gvisor.dev/gvisor/pkg/rand" "gvisor.dev/gvisor/pkg/sentry/fsimpl/cgroup2fs" "gvisor.dev/gvisor/pkg/sentry/fsimpl/host" @@ -50,7 +50,7 @@ var sharedPlatform platform.Platform var platformErr error // Systrap owns process-lifetime memory and stub pools. Each Sandbox owns a -// Kernel; its fresh guest processes share that Kernel and platform. +// Kernel; its resident guest processes share that Kernel and platform. func newKernel(p *observedPlatform) (*kernel.Kernel, error) { cpuid.Initialize() seccheck.Initialize() @@ -77,12 +77,22 @@ func newKernel(p *observedPlatform) (*kernel.Kernel, error) { p.Platform = sharedPlatform k := &kernel.Kernel{Platform: p} - memoryFD, err := memutil.CreateMemFD("llar-runtime-memory", 0) + // A disk file lets idle processes release resident pages without losing + // their virtual addresses. Unlink it immediately; Kernel owns its lifetime. + memoryFile, err := os.CreateTemp("", "llar-runtime-memory-*") if err != nil { return nil, err } - memoryFile := os.NewFile(uintptr(memoryFD), "llar-runtime-memory") - mf, err := pgalloc.NewMemoryFile(memoryFile, pgalloc.MemoryFileOpts{}) + if err := os.Remove(memoryFile.Name()); err != nil { + memoryFile.Close() + return nil, err + } + var stat unix.Statfs_t + if err := unix.Fstatfs(int(memoryFile.Fd()), &stat); err != nil { + memoryFile.Close() + return nil, err + } + mf, err := pgalloc.NewMemoryFile(memoryFile, pgalloc.MemoryFileOpts{DiskBackedFile: stat.Type != unix.TMPFS_MAGIC && stat.Type != unix.RAMFS_MAGIC, DecommitOnDestroy: true}) if err != nil { memoryFile.Close() return nil, err diff --git a/sentry/fs_linux.go b/sentry/fs_linux.go index 5d600a5..f959881 100644 --- a/sentry/fs_linux.go +++ b/sentry/fs_linux.go @@ -245,11 +245,11 @@ func mountFilesystem(ctx context.Context, k *kernel.Kernel, mounts []mount) (_ * return mntns, nil } -func importDescriptors(k *kernel.Kernel, imageFD int) (*kernel.FDTable, error) { +func importDescriptors(k *kernel.Kernel, imageFD, controlFD int) (*kernel.FDTable, error) { ctx := k.SupervisorContext() table := k.NewFDTable() files := make(map[int]*fd.FD) - for n, descriptor := range []int{int(os.Stdin.Fd()), int(os.Stdout.Fd()), int(os.Stderr.Fd()), imageFD} { + for n, descriptor := range []int{int(os.Stdin.Fd()), int(os.Stdout.Fd()), int(os.Stderr.Fd()), imageFD, controlFD} { dup, err := unix.Dup(descriptor) if err != nil { table.DecRef(ctx) @@ -263,5 +263,11 @@ func importDescriptors(k *kernel.Kernel, imageFD int) (*kernel.FDTable, error) { table.DecRef(ctx) return nil, err } + for _, descriptor := range []int32{3, 4} { + if err := table.SetFlags(ctx, descriptor, kernel.FDFlags{CloseOnExec: true}); err != nil { + table.DecRef(ctx) + return nil, err + } + } return table, nil } diff --git a/sentry/main_linux.go b/sentry/main_linux.go index d27e0cc..55060b7 100644 --- a/sentry/main_linux.go +++ b/sentry/main_linux.go @@ -79,17 +79,44 @@ func CloseSandbox(handle C.uintptr_t, message *C.char, capacity C.size_t) C.int } s.mu.Lock() s.closed = true - s.kernel.Kill(linux.WaitStatusExit(1)) + processes := make([]*guestProcess, 0, len(s.processes)) + for _, p := range s.processes { + processes = append(processes, p) + } s.mu.Unlock() + var closeErr error + for _, p := range processes { + closeErr = errors.Join(closeErr, p.close()) + } s.runs.Wait() + s.kernel.Kill(linux.WaitStatusExit(1)) s.kernel.WaitExited() s.dog.Stop() s.kernel.Release() - return 0 + return reportError(closeErr, message, capacity) } //export RunSandbox func RunSandbox(handle C.uintptr_t, config *C.char, imageFD C.int, mainPC, entryPC, owner C.uintptr_t, callback C.inspect_fn, message *C.char, capacity C.size_t) (code C.int) { + runtime.LockOSThread() + defer runtime.UnlockOSThread() + report := func(err error) { code = reportError(err, message, capacity) } + defer func() { + if v := recover(); v != nil { + report(fmt.Errorf("Sentry process operation panicked: %v", v)) + } + }() + var startup struct { + Operation string `json:"operation"` + Process uint64 `json:"process"` + Guest string `json:"guest"` + Mounts []mount `json:"mounts"` + Env []string `json:"env"` + } + if err := json.Unmarshal([]byte(C.GoString(config)), &startup); err != nil { + report(fmt.Errorf("Sentry configuration: %w", err)) + return + } kernels.Lock() s := kernels.live[uintptr(handle)] if s != nil { @@ -97,31 +124,39 @@ func RunSandbox(handle C.uintptr_t, config *C.char, imageFD C.int, mainPC, entry } kernels.Unlock() if s == nil { - return reportError(errors.New("sandbox is closed"), message, capacity) + if startup.Operation != "close" { + report(errors.New("sandbox is closed")) + } + return } defer s.runs.Done() - runtime.LockOSThread() - defer runtime.UnlockOSThread() - report := func(err error) { - code = reportError(err, message, capacity) - } - defer func() { - if v := recover(); v != nil { - report(fmt.Errorf("Sentry startup panicked: %v", v)) + if startup.Operation == "close" || startup.Operation == "run" { + s.mu.Lock() + process := s.processes[startup.Process] + if startup.Operation == "close" { + delete(s.processes, startup.Process) } - }() - var startup struct { - Guest string `json:"guest"` - Mounts []mount `json:"mounts"` - Env []string `json:"env"` + closed := s.closed + s.mu.Unlock() + if startup.Operation == "close" { + if process != nil { + report(process.close()) + } + return + } + if closed || process == nil { + report(errors.New("process is closed or missing")) + return + } + report(process.run(int(imageFD))) + return } - if err := json.Unmarshal([]byte(C.GoString(config)), &startup); err != nil { - report(fmt.Errorf("Sentry startup configuration: %w", err)) - return code + if startup.Operation != "create" { + report(errors.New("invalid process operation")) + return } - var inspectionMu sync.Mutex - var inspectionErr error - err := s.run(startup.Mounts, startup.Guest, startup.Env, int(imageFD), uintptr(mainPC), uintptr(entryPC), func(ctx gcontext.Context, ac *arch.Context64) error { + process := &guestProcess{} + process.inspect = func(ctx gcontext.Context, ac *arch.Context64) error { if callback == nil { return nil } @@ -154,11 +189,11 @@ func RunSandbox(handle C.uintptr_t, config *C.char, imageFD C.int, mainPC, entry } } if err != nil { - inspectionMu.Lock() - if inspectionErr == nil { - inspectionErr = err + process.inspectionMu.Lock() + if process.inspectionErr == nil { + process.inspectionErr = err } - inspectionMu.Unlock() + process.inspectionMu.Unlock() _ = task.Kernel().SendContainerSignal(task.ContainerID(), &linux.SignalInfo{Signo: int32(linux.SIGKILL)}) return err } @@ -169,14 +204,9 @@ func RunSandbox(handle C.uintptr_t, config *C.char, imageFD C.int, mainPC, entry setSyscall(ac, uint64(event.number), args) ac.SyscallSaveOrig() return nil - }) - if inspectionErr != nil { - err = inspectionErr - } - if err != nil { - report(err) } - return code + report(s.start(startup.Process, process, startup.Mounts, startup.Guest, startup.Env, int(imageFD), uintptr(mainPC), uintptr(entryPC))) + return } func main() {} diff --git a/sentry/platform_linux.go b/sentry/platform_linux.go index 27cb084..58b292b 100644 --- a/sentry/platform_linux.go +++ b/sentry/platform_linux.go @@ -27,17 +27,39 @@ func (p *observedPlatform) NewContext(ctx context.Context) platform.Context { process := p.processes[task.ContainerID()] process.tasks.Add(1) p.mu.Unlock() - return &observedContext{Context: p.Platform.NewContext(ctx), process: process} + c := &observedContext{Context: p.Platform.NewContext(ctx), process: process, task: task} + // NewTask holds its signal lock here and has not assigned task.p yet. + // Only register the context; pause requests its stop outside taskMu. + process.taskMu.Lock() + process.contexts[c] = false + process.taskMu.Unlock() + return c } type observedContext struct { platform.Context process *guestProcess + task *kernel.Task + mu sync.Mutex // Serializes stop requests with platform context release. } func (c *observedContext) Release() { + c.mu.Lock() + p := c.process + p.taskMu.Lock() + stopped := p.contexts[c] + p.taskMu.Unlock() + // A task can finish its exit path without visiting doStop again. Balance + // any request that raced with exit before releasing the platform context. + if stopped { + c.task.EndExternalStop() + } c.Context.Release() - c.process.tasks.Done() + p.taskMu.Lock() + delete(p.contexts, c) + p.taskMu.Unlock() + c.mu.Unlock() + p.tasks.Done() } func (c *observedContext) Switch(ctx context.Context, mm platform.MemoryManager, ac *arch.Context64, cpu int32) (*linux.SignalInfo, hostarch.AccessType, error) { diff --git a/sentry/process_memory_linux.go b/sentry/process_memory_linux.go new file mode 100644 index 0000000..80c59dc --- /dev/null +++ b/sentry/process_memory_linux.go @@ -0,0 +1,129 @@ +//go:build linux && (amd64 || arm64) && cgo + +package main + +import ( + "errors" + "fmt" + "reflect" + + "golang.org/x/sys/unix" + "gvisor.dev/gvisor/pkg/abi/linux" + "gvisor.dev/gvisor/pkg/context" + "gvisor.dev/gvisor/pkg/hostarch" + "gvisor.dev/gvisor/pkg/sentry/kernel" + "gvisor.dev/gvisor/pkg/sentry/memmap" + "gvisor.dev/gvisor/pkg/sentry/mm" + "gvisor.dev/gvisor/pkg/sentry/pgalloc" +) + +// evict is called with p.mu held after every task in this process has stopped. +// It retains VMAs, PMAs and thread state; the ordinary fault path reloads pages. +func (p *guestProcess) evict() error { + ctx := p.kernel.SupervisorContext() + managers := make(map[*mm.MemoryManager]struct{}) + for _, task := range p.kernel.RootPIDNamespace().Tasks() { + if task.ContainerID() != p.id { + continue + } + task.WithMuLocked(func(task *kernel.Task) { + m := task.MemoryManager() + if m == nil { + return + } + if _, exists := managers[m]; exists { + return + } + if m.IncUsers() { + managers[m] = struct{}{} + } + }) + } + defer func() { + for m := range managers { + m.DecUsers(ctx) + } + }() + var err error + for m := range managers { + err = errors.Join(err, evictPrivateMemory(ctx, m, p.kernel.MemoryFile())) + } + return err +} + +// gVisor d1e35511e5a4 keeps its resident mappings in MemoryManager.pmas. +// Pin would instantiate untouched reservations and break COW, so read the +// existing set through its ExportSlice method while holding its own activeMu. +// Reflection avoids duplicating the generated segment-tree layout. This adapter +// is local to Sentry; gVisor's source and public API remain unchanged. +func evictPrivateMemory(ctx context.Context, m *mm.MemoryManager, mf *pgalloc.MemoryFile) (err error) { + defer func() { + if v := recover(); v != nil { + err = fmt.Errorf("gVisor private mapping metadata: %v", v) + } + }() + if !mf.IsDiskBacked() { + return errors.New("guest memory backing is not a disk file") + } + value := reflect.ValueOf(m).Elem() + lockField := value.FieldByName("activeMu") + lock := reflect.NewAt(lockField.Type(), lockField.Addr().UnsafePointer()).Interface().(interface { + Lock() + Unlock() + }) + // Keep kernel-side memory I/O and mapping invalidation out of this operation, + // including asynchronous I/O that does not execute a guest task. + lock.Lock() + defer lock.Unlock() + set := value.FieldByName("pmas") + flat := reflect.NewAt(set.Type(), set.Addr().UnsafePointer()).MethodByName("ExportSlice").Call(nil)[0] + var pins []mm.PinnedRange + defer func() { mm.Unpin(pins) }() + for i := 0; i < flat.Len(); i++ { + segment := flat.Index(i) + page := segment.FieldByName("Value") + if !page.FieldByName("private").Bool() || page.FieldByName("needCOW").Bool() { + continue + } + fileField := page.FieldByName("file") + file := reflect.NewAt(fileField.Type(), fileField.Addr().UnsafePointer()).Elem().Interface().(memmap.File) + if file != mf { + return errors.New("private mapping has an unexpected backing file") + } + pin := mm.PinnedRange{Source: hostarch.AddrRange{Start: hostarch.Addr(segment.FieldByName("Start").Uint()), End: hostarch.Addr(segment.FieldByName("End").Uint())}, File: file, Offset: page.FieldByName("off").Uint()} + file.IncRef(pin.FileRange(), pgalloc.MemoryCgroupIDFromContext(ctx)) + pins = append(pins, pin) + } + // Complete writeback before removing any mappings. Decommit is deliberately + // not used: it punches holes and destroys the stored guest contents. + for _, pin := range pins { + blocks, e := mf.MapInternal(pin.FileRange(), hostarch.ReadWrite) + if e != nil { + return e + } + for !blocks.IsEmpty() { + if e := unix.Msync(blocks.Head().ToSlice(), unix.MS_SYNC); e != nil { + return e + } + blocks = blocks.Tail() + } + } + for _, pin := range pins { + m.AddressSpace().Unmap(pin.Source.Start, uint64(pin.Source.Length())) + blocks, e := mf.MapInternal(pin.FileRange(), hostarch.ReadWrite) + if e != nil { + return e + } + for !blocks.IsEmpty() { + if e := unix.Madvise(blocks.Head().ToSlice(), unix.MADV_DONTNEED); e != nil { + return e + } + blocks = blocks.Tail() + } + fr := pin.FileRange() + if e := unix.Fadvise(mf.FD(), int64(fr.Start), int64(fr.Length()), linux.POSIX_FADV_DONTNEED); e != nil { + return e + } + } + return nil +} diff --git a/sentry/run_linux.go b/sentry/run_linux.go index a1b05b8..a1c2a17 100644 --- a/sentry/run_linux.go +++ b/sentry/run_linux.go @@ -1,18 +1,24 @@ -//go:build linux && (arm64 || amd64) && cgo +//go:build linux && (amd64 || arm64) && cgo package main import ( "errors" "fmt" + "io" + "os" "path" "strconv" "strings" "sync" + "time" + "golang.org/x/sys/unix" "gvisor.dev/gvisor/pkg/abi/linux" "gvisor.dev/gvisor/pkg/context" "gvisor.dev/gvisor/pkg/hostarch" + "gvisor.dev/gvisor/pkg/log" + "gvisor.dev/gvisor/pkg/sentry/fsimpl/host" "gvisor.dev/gvisor/pkg/sentry/kernel" "gvisor.dev/gvisor/pkg/sentry/kernel/auth" "gvisor.dev/gvisor/pkg/sentry/limits" @@ -22,21 +28,37 @@ import ( ) type sentryKernel struct { - mu sync.Mutex // Serializes process creation/start against Close. - kernel *kernel.Kernel - platform *observedPlatform - dog *watchdog.Watchdog - runs sync.WaitGroup - closed bool + mu sync.Mutex + kernel *kernel.Kernel + platform *observedPlatform + dog *watchdog.Watchdog + runs sync.WaitGroup + closed bool + processes map[uint64]*guestProcess } type guestProcess struct { - inspect inspector - tasks sync.WaitGroup + inspect inspector + tasks sync.WaitGroup + taskMu sync.Mutex // Protects registration and external-stop ownership. + contexts map[*observedContext]bool + kernel *kernel.Kernel + id string + tg *kernel.ThreadGroup + control *os.File + mu sync.Mutex // Serializes commands, pauses and idle eviction. + paused bool + idleTimer *time.Timer + closeOnce sync.Once + stopping chan struct{} + exited chan struct{} + done chan struct{} + exitErr error + closeErr error + inspectionMu sync.Mutex + inspectionErr error } -// Kernel supervisor contexts fall back to the first task's mount namespace. -// A new Run has no mount namespace until its own root has been mounted. type mountContext struct{ context.Context } func (c mountContext) Value(key any) any { @@ -46,12 +68,12 @@ func (c mountContext) Value(key any) any { return c.Context.Value(key) } -func (s *sentryKernel) run(mounts []mount, executable string, env []string, imageFD int, mainPC, entryPC uintptr, inspect inspector) (err error) { +func (s *sentryKernel) start(processID uint64, p *guestProcess, mounts []mount, executable string, env []string, imageFD int, mainPC, entryPC uintptr) (err error) { if !path.IsAbs(executable) || strings.ContainsRune(executable, 0) { return errors.New("guest executable must be an absolute path without NUL") } for _, entry := range env { - if strings.IndexByte(entry, 0) >= 0 { + if strings.ContainsRune(entry, 0) { return errors.New("guest environment contains NUL") } } @@ -59,86 +81,327 @@ func (s *sentryKernel) run(mounts []mount, executable string, env []string, imag if err != nil { return err } - defer closeMounts(mounts) + defer func() { + if err != nil { + closeMounts(mounts) + } + }() if err := prepareMounts(mounts); err != nil { return err } - + pair, err := unix.Socketpair(unix.AF_UNIX, unix.SOCK_STREAM|unix.SOCK_CLOEXEC|unix.SOCK_NONBLOCK, 0) + if err != nil { + return err + } + p.control = os.NewFile(uintptr(pair[0]), "sandbox-control") + defer unix.Close(pair[1]) + defer func() { + if err != nil { + p.control.Close() + } + }() + // The guest sees a blocking descriptor. host.NewFD makes its host backing + // nonblocking and implements guest blocking through Sentry's wait queues. + if err := unix.SetNonblock(pair[1], false); err != nil { + return err + } k := s.kernel + p.kernel, p.id = k, strconv.FormatUint(processID, 10) + p.stopping, p.exited, p.done = make(chan struct{}), make(chan struct{}), make(chan struct{}) + p.contexts = make(map[*observedContext]bool) pidns := k.RootPIDNamespace().NewChild(k.SupervisorContext(), k, k.RootUserNamespace()) defer pidns.DecRef(k.SupervisorContext()) - id := strconv.FormatUint(pidns.ID(), 10) - process := &guestProcess{inspect: inspect} - s.platform.mu.Lock() - s.platform.processes[id] = process - s.platform.mu.Unlock() - defer func() { - // Platform contexts are released after their task has dropped its MM, - // FDs and mounts. Include forked children, even after they are reaped. - process.tasks.Wait() - err = errors.Join(err, releaseSyscallMemory(k, id)) - s.platform.mu.Lock() - delete(s.platform.processes, id) - s.platform.mu.Unlock() - }() ls, err := limits.NewLinuxLimitSet() if err != nil { return err } - args := kernel.CreateProcessArgs{ - Filename: executable, Argv: []string{executable}, Envv: env, - WorkingDirectory: "/", - Credentials: auth.NewUserCredentials(1000, 1000, nil, &auth.TaskCapabilities{}, k.RootUserNamespace()), - Umask: 0022, Limits: ls, - MaxSymlinkTraversals: linux.MaxSymlinkTraversals, - UTSNamespace: k.RootUTSNamespace(), IPCNamespace: k.RootIPCNamespace(), - PIDNamespace: pidns, ContainerID: id, - } + args := kernel.CreateProcessArgs{Filename: executable, Argv: []string{executable}, Envv: env, WorkingDirectory: "/", Credentials: auth.NewUserCredentials(1000, 1000, nil, &auth.TaskCapabilities{}, k.RootUserNamespace()), Umask: 0022, Limits: ls, MaxSymlinkTraversals: linux.MaxSymlinkTraversals, UTSNamespace: k.RootUTSNamespace(), IPCNamespace: k.RootIPCNamespace(), PIDNamespace: pidns, ContainerID: p.id} ctx := mountContext{args.NewContext(k)} mntns, err := mountFilesystem(ctx, k, mounts) if err != nil { return err } defer mntns.DecRef(ctx) - - fdt, err := importDescriptors(k, imageFD) + table, err := importDescriptors(k, imageFD, pair[1]) if err != nil { return err } - defer fdt.DecRef(ctx) - args.FDTable, args.MountNamespace = fdt, mntns + defer table.DecRef(ctx) + args.FDTable, args.MountNamespace = table, mntns s.mu.Lock() if s.closed { s.mu.Unlock() return errors.New("sandbox is closed") } - // CreateProcess takes ownership of one mount-namespace reference. The local - // reference above remains valid for cleanup on both success and failure. + if _, exists := s.processes[processID]; exists { + s.mu.Unlock() + return errors.New("process already exists") + } + s.platform.mu.Lock() + s.platform.processes[p.id] = p + s.platform.mu.Unlock() mntns.IncRef() tg, _, err := k.CreateProcess(args) if err != nil { + s.platform.mu.Lock() + delete(s.platform.processes, p.id) + s.platform.mu.Unlock() s.mu.Unlock() return fmt.Errorf("loading guest: %w", err) } - - // Patch only the guest's private ELF mapping, before any guest task runs. - // Go initializes its runtime and packages, then main.main branches to entryPC. + p.tg = tg _, err = tg.Leader().MemoryManager().CopyOut(ctx, hostarch.Addr(mainPC), jump, usermem.IOOpts{IgnorePermissions: true}) if err != nil { - // Even an unstarted task must enter the task loop to release its refs. _ = tg.SendSignal(&linux.SignalInfo{Signo: int32(linux.SIGKILL)}) } k.StartProcess(tg) - s.mu.Unlock() - tg.WaitExited() - process.tasks.Wait() if err != nil { + s.mu.Unlock() + tg.WaitExited() + p.tasks.Wait() + s.platform.mu.Lock() + delete(s.platform.processes, p.id) + s.platform.mu.Unlock() return fmt.Errorf("installing guest entry: %w", err) } + if s.processes == nil { + s.processes = make(map[uint64]*guestProcess) + } + s.processes[processID] = p + go func() { + tg.WaitExited() + p.tasks.Wait() + p.control.Close() + status := tg.ExitStatus() + if !status.Exited() || status.ExitStatus() != 0 { + p.exitErr = fmt.Errorf("application exited with %s", status) + } + close(p.exited) + p.mu.Lock() + if p.idleTimer != nil { + p.idleTimer.Stop() + p.idleTimer = nil + } + p.closeErr = releaseSyscallMemory(k, p.id) + closeMounts(mounts) + s.platform.mu.Lock() + delete(s.platform.processes, p.id) + s.platform.mu.Unlock() + p.mu.Unlock() + close(p.done) + }() + s.mu.Unlock() + return nil +} - status := tg.ExitStatus() - if !status.Exited() || status.ExitStatus() != 0 { - return fmt.Errorf("application exited with %s", status) +func (p *guestProcess) run(imageFD int) (runErr error) { + p.mu.Lock() + defer p.mu.Unlock() + defer func() { + // An inspection failure terminates the guest and closes its control + // socket. Preserve that cause instead of reporting the resulting EOF. + p.inspectionMu.Lock() + if p.inspectionErr != nil { + runErr = p.inspectionErr + } + p.inspectionMu.Unlock() + }() + select { + case <-p.stopping: + return errors.New("process is closed") + case <-p.exited: + return errors.Join(errors.New("process has exited"), p.exitErr) + default: + } + if p.idleTimer != nil { + p.idleTimer.Stop() + p.idleTimer = nil } + if p.paused { + if err := p.installImage(imageFD); err != nil { + return err + } + p.resume() + } + command := [1]byte{1} + _, err := p.control.Write(command[:]) + if err == nil { + _, err = io.ReadFull(p.control, command[:]) + } + if err == nil && command[0] != 1 { + err = errors.New("invalid guest completion message") + } + if err != nil { + _ = p.kernel.SendContainerSignal(p.id, &linux.SignalInfo{Signo: int32(linux.SIGKILL)}) + <-p.exited + return errors.Join(err, p.exitErr) + } + if err := p.pause(); err != nil { + return err + } + p.inspectionMu.Lock() + inspectionErr := p.inspectionErr + p.inspectionMu.Unlock() + if inspectionErr != nil { + return inspectionErr + } + var timer *time.Timer + timer = time.AfterFunc(10*time.Second, func() { + p.mu.Lock() + defer p.mu.Unlock() + if p.idleTimer != timer { + return + } + p.idleTimer = nil + select { + case <-p.stopping: + return + case <-p.exited: + return + default: + } + if err := p.evict(); err != nil { + log.Warningf("Idle page eviction for process %s failed: %v", p.id, err) + } + }) + p.idleTimer = timer return nil } + +func (p *guestProcess) installImage(imageFD int) error { + ctx := p.kernel.SupervisorContext() + duplicate, err := unix.Dup(imageFD) + if err != nil { + return err + } + file, err := host.NewFD(ctx, p.kernel.HostMount(), duplicate, &host.NewFDOptions{Savable: true, VirtualOwner: true, UID: 1000, GID: 1000}) + if err != nil { + unix.Close(duplicate) + return err + } + defer file.DecRef(ctx) + var table *kernel.FDTable + p.tg.Leader().WithMuLocked(func(task *kernel.Task) { + table = task.FDTable() + if table != nil { + table.IncRef() + } + }) + if table == nil { + return errors.New("process descriptor table has been released") + } + defer table.DecRef(ctx) + old, err := table.NewFDAt(ctx, 3, file, kernel.FDFlags{CloseOnExec: true}) + if old != nil { + old.DecRef(ctx) + } + return err +} + +// pause is called with p.mu held. Newly registered contexts join the same stop +// round. For example, a clone already executing when pause starts may register +// its child after the first pass; its parent cannot acknowledge until it does. +func (p *guestProcess) pause() (err error) { + defer func() { + if err != nil { + p.resume() + } + }() + ticker := time.NewTicker(time.Millisecond) + defer ticker.Stop() + for { + p.taskMu.Lock() + var pending []*observedContext + for c, stopped := range p.contexts { + if !stopped { + pending = append(pending, c) + } + } + p.taskMu.Unlock() + for _, c := range pending { + c.mu.Lock() + p.taskMu.Lock() + _, live := p.contexts[c] + p.taskMu.Unlock() + if live { + // Never hold taskMu across gVisor calls: NewContext registers + // while holding the signal lock that BeginExternalStop needs. + c.task.BeginExternalStop() + p.taskMu.Lock() + p.contexts[c] = true + p.taskMu.Unlock() + } + c.mu.Unlock() + } + p.taskMu.Lock() + all := len(p.contexts) > 0 + for c, stopped := range p.contexts { + if !stopped || c.task.TaskGoroutineState() != kernel.TaskGoroutineStopped { + all = false + break + } + } + // Check membership and acknowledgements together. Once all tasks are + // stopped, none can create another task until this process resumes. + p.taskMu.Unlock() + if all { + p.paused = true + return nil + } + select { + case <-p.stopping: + return errors.New("process is closing") + case <-p.exited: + return errors.Join(errors.New("process exited before pausing"), p.exitErr) + case <-ticker.C: + } + } +} + +// resume releases only this process's external stops. p.mu must be held, so +// pause cannot add requests while this pass releases them. +func (p *guestProcess) resume() { + p.taskMu.Lock() + var stopped []*observedContext + for c, held := range p.contexts { + if held { + stopped = append(stopped, c) + } + } + p.taskMu.Unlock() + for _, c := range stopped { + c.mu.Lock() + p.taskMu.Lock() + held := p.contexts[c] + if held { + p.contexts[c] = false + } + p.taskMu.Unlock() + if held { + c.task.EndExternalStop() + } + c.mu.Unlock() + } + p.paused = false +} + +func (p *guestProcess) close() error { + p.closeOnce.Do(func() { + close(p.stopping) + // Closing a pollable control socket wakes an active run. Take mu only + // afterwards so Close can interrupt a callback that never returns. + p.control.Close() + p.mu.Lock() + if p.idleTimer != nil { + p.idleTimer.Stop() + p.idleTimer = nil + } + _ = p.kernel.SendContainerSignal(p.id, &linux.SignalInfo{Signo: int32(linux.SIGKILL)}) + // SIGKILL does not release external stops. Queue it first, then let + // stopped tasks run their exit paths before waiting for cleanup. + p.resume() + p.mu.Unlock() + <-p.done + }) + return p.closeErr +} diff --git a/sentry/sandbox.h b/sentry/sandbox.h index 1c98608..d49d795 100644 --- a/sentry/sandbox.h +++ b/sentry/sandbox.h @@ -20,7 +20,8 @@ struct syscall_event { typedef void (*inspect_fn)(uintptr_t, struct syscall_event *); typedef int (*create_sentry_fn)(uintptr_t *, char *, size_t); typedef int (*close_sentry_fn)(uintptr_t, char *, size_t); -// Each run receives a Kernel handle and startup JSON (guest, mounts and env). +// Dispatches create/run/close for a process ID within a Kernel. +// Creation supplies guest/mount/env configuration; each run supplies a fresh image fd. typedef int (*run_sentry_fn)(uintptr_t, char *, int, uintptr_t, uintptr_t, uintptr_t, inspect_fn, char *, size_t); diff --git a/testdata/main.go b/testdata/main.go index 6b2ee9d..d0c920e 100644 --- a/testdata/main.go +++ b/testdata/main.go @@ -38,6 +38,9 @@ func main() { } func run() error { + if err := checkProcessLifecycle(); err != nil { + return fmt.Errorf("persistent process: %w", err) + } if err := checkKernelLifecycle(); err != nil { return fmt.Errorf("kernel lifecycle: %w", err) } diff --git a/testdata/process.go b/testdata/process.go new file mode 100644 index 0000000..4d50d70 --- /dev/null +++ b/testdata/process.go @@ -0,0 +1,200 @@ +//go:build linux && (amd64 || arm64) && cgo + +package main + +import ( + "fmt" + "mime" + "os" + "sync/atomic" + "time" + + "github.com/xgo-dev/sandbox" + "golang.org/x/sys/unix" +) + +var residentNumber int +var residentPointer *int +var residentBytes []byte +var residentStop chan struct{} + +func checkProcessLifecycle() error { + const heartbeat = 0xffffffd0 + var ticks atomic.Int64 + p := sandbox.NewProcess(sandbox.ProcessOptions{Inspect: func(call *sandbox.Syscall) { + if call.Number == heartbeat { + ticks.Add(1) + call.Number = unix.SYS_GETPID + } + }}) + defer p.Close() + x := 40 + if err := p.Run(func() { + residentNumber = 7 + residentPointer = &x + x++ + residentBytes = make([]byte, 64<<20) + for i := range residentBytes { + residentBytes[i] = byte(i*17 + 29) + } + residentStop = make(chan struct{}) + if err := mime.AddExtensionType(".sandbox-resident", "application/x-sandbox-resident"); err != nil { + panic(err) + } + go func() { + ticker := time.NewTicker(10 * time.Millisecond) + defer ticker.Stop() + for { + select { + case <-residentStop: + return + case <-ticker.C: + unix.RawSyscall(heartbeat, 0, 0, 0) + } + } + }() + // Publish one observation before the callback completes. + unix.RawSyscall(heartbeat, 0, 0, 0) + }); err != nil { + return fmt.Errorf("persistent initial run: %w", err) + } + if x != 41 || residentNumber != 0 || residentPointer != nil || residentBytes != nil { + return fmt.Errorf("persistent first writeback/global isolation: x=%d global=%d", x, residentNumber) + } + paused := ticks.Load() + time.Sleep(50 * time.Millisecond) + if ticks.Load() != paused { + return fmt.Errorf("guest kept executing after Run returned") + } + x = 50 + if err := p.Run(func() { + check(residentNumber == 7 && residentPointer == &x && *residentPointer == 50, "resident global alias") + check(mime.TypeByExtension(".sandbox-resident") == "application/x-sandbox-resident", "native MIME state") + x++ + residentNumber++ + }); err != nil { + return fmt.Errorf("persistent second run: %w", err) + } + if x != 51 { + return fmt.Errorf("persistent second writeback: %d", x) + } + q := sandbox.NewProcess() + defer q.Close() + independent := false + if err := q.Run(func() { independent = residentNumber == 0 && residentBytes == nil && residentPointer == nil }); err != nil { + return err + } + if !independent { + return fmt.Errorf("two processes shared native globals") + } + paused = ticks.Load() + // Exercise the actual ten-second idle policy, then fault guest pages back. + time.Sleep(11 * time.Second) + if ticks.Load() != paused { + return fmt.Errorf("idle eviction resumed the guest") + } + if err := q.Run(func() { check(residentNumber == 0, "other process changed during eviction") }); err != nil { + return err + } + if err := p.Run(func() { + check(residentNumber == 8 && residentPointer == &x, "resident state after eviction") + for i, v := range residentBytes { + check(v == byte(i*17+29), "resident bytes after eviction") + } + check(mime.TypeByExtension(".sandbox-resident") == "application/x-sandbox-resident", "native MIME after eviction") + close(residentStop) + x++ + }); err != nil { + return fmt.Errorf("persistent run after idle: %w", err) + } + if x != 52 { + return fmt.Errorf("persistent idle writeback: %d", x) + } + if err := p.Close(); err != nil { + return err + } + if err := p.Close(); err != nil { + return err + } + if err := p.Run(staticCall); err == nil { + return fmt.Errorf("closed process accepted Run") + } + if err := q.Run(staticCall); err != nil { + return fmt.Errorf("closing one process affected another: %w", err) + } + + // Concurrent submissions to the same process must execute serially. + serial := sandbox.NewProcess() + defer serial.Close() + results := make(chan error, 2) + first, second := 0, 0 + go func() { results <- serial.Run(func() { residentNumber++; first = residentNumber }) }() + go func() { results <- serial.Run(func() { residentNumber++; second = residentNumber }) }() + for range 2 { + if err := <-results; err != nil { + return err + } + } + if first+second != 3 || first == second { + return fmt.Errorf("same-process calls overlapped or restarted: %d %d", first, second) + } + + const waiting = 0xffffffd1 + ready := make(chan struct{}) + active := sandbox.NewProcess(sandbox.ProcessOptions{Inspect: func(call *sandbox.Syscall) { + if call.Number == waiting { + close(ready) + call.Number = unix.SYS_GETPID + } + }}) + defer active.Close() + original := x + go func() { + results <- active.Run(func() { + x = 999 + unix.RawSyscall(waiting, 0, 0, 0) + for { + time.Sleep(time.Second) + } + }) + }() + select { + case <-ready: + case err := <-results: + return fmt.Errorf("active process failed before Close: %v", err) + case <-time.After(20 * time.Second): + return fmt.Errorf("active process did not start") + } + if err := active.Close(); err != nil { + return err + } + if err := <-results; err == nil { + return fmt.Errorf("terminated process Run succeeded") + } + if x != original { + return fmt.Errorf("terminated process committed captures") + } + if err := q.Run(staticCall); err != nil { + return err + } + + // Inherited environment is captured at process startup, while explicit empty + // environment remains empty. This process's init sees the supplied value. + configured := sandbox.NewProcess(sandbox.ProcessOptions{Env: []string{"SANDBOX_ENV_TEST=resident"}}) + defer configured.Close() + env := "" + if err := configured.Run(func() { env = environmentAtInit; _ = os.Setenv("SANDBOX_ENV_TEST", "changed") }); err != nil { + return err + } + if env != "resident" { + return fmt.Errorf("process startup environment: %q", env) + } + if err := configured.Run(func() { env = os.Getenv("SANDBOX_ENV_TEST") }); err != nil { + return err + } + if env != "changed" { + return fmt.Errorf("process environment did not persist: %q", env) + } + fmt.Println("PASS persistent globals, captured aliases, pause, idle recovery, process isolation, serialized Run, environment persistence and Close") + return nil +} diff --git a/unsupported.go b/unsupported.go index b526f27..7400bb3 100644 --- a/unsupported.go +++ b/unsupported.go @@ -4,6 +4,14 @@ package sandbox import "errors" +type processState struct{} + +func (*Process) Run(func()) error { + return errors.New("sandbox requires Linux, amd64 or arm64, and cgo") +} + +func (*Process) Close() error { return nil } + func (*Sandbox) Run(func()) error { return errors.New("sandbox requires Linux, amd64 or arm64, and cgo") }