Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 28 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,9 +12,33 @@ func main() {
}
```

Guest startup is internal to the library. A new guest executes Go runtime and package initialization, then enters the imported closure without executing the application's `main`. Initialization side effects must be appropriate there. Captured objects must be exclusively owned until `Run` returns. The function must finish its own goroutines before returning. Concurrent and nested host calls are supported when their captured graphs are independent.
Guest startup is internal to the library. A new guest executes Go runtime and package initialization, then enters the imported closure without executing the application's `main`. Initialization side effects must be appropriate there. Captured objects must be exclusively owned until `Run` returns. Goroutines must stop accessing captured objects before the callback returns; a persistent Process may retain goroutines that only use guest-owned state. Concurrent and nested host calls are supported when their captured graphs are independent.

`sandbox.Run` uses one global default Sandbox, initialized on first use and retained for the host process lifetime. Each explicit `Sandbox` likewise starts its Kernel on its first `Run` and reuses it until `Close`. Every call creates a fresh guest process, PID namespace, filesystem namespace, FD table and snapshot, and transfers the complete type and object graph. Guest processes are not pooled. `Close` rejects new calls, terminates active guests and waits for their calls to finish; repeated `Close` calls are safe. Do not call it from that Sandbox's inspector. Do not copy a used Sandbox or change its fields while calls are active; `Library` must remain unchanged after first use.
`sandbox.Run` uses one global default Sandbox, initialized on first use and retained for the host process lifetime. Each explicit `Sandbox` likewise reuses its Kernel until `Close`. These one-shot entry points create and close a fresh guest for each call. `Sandbox.Close` rejects new calls, terminates its guests and waits for active calls to finish. Do not call Close from the corresponding inspector, copy a used Sandbox, or change its configuration while calls are active; `Library` remains fixed after first use.

## Persistent processes

Use one Process when callbacks need the same native package globals, open guest files or filesystem state:

```go
p := sandbox.NewProcess(sandbox.ProcessOptions{
Env: os.Environ(),
})
defer p.Close()

if err := p.Run(func() { f.OnRequire(proj, deps) }); err != nil {
return err
}
if err := p.Run(func() { f.OnBuild(ctx) }); err != nil {
return err
}
```

`NewProcess` accepts zero or one `ProcessOptions` value. It uses the default shared Kernel and starts its guest on the first Run. Options contain `Mounts`, `Env` and `Inspect`; nil Env inherits the host environment at startup. The guest's environment and mounts then persist. Each successful Run writes captured changes back, pauses this process's guest tasks and leaves native global state in the guest. Calls on one Process execute serially; independent processes may run concurrently. Existing object IDs are retained across calls, including when a guest global keeps a pointer to a captured object. Close terminates the process and any active call, waits for cleanup, and is idempotent. Failed execution or a failed transfer closes the process; it is never silently restarted.

After ten seconds idle, Sentry flushes eligible private pages and asks the host to reclaim their resident storage. The process, virtual mappings and kernel resources stay alive; subsequent accesses fault pages back at the same addresses. The application MemoryFile is backed by an unlinked temporary file. Place the host temporary directory on a disk filesystem for reclamation; a tmpfs-backed file cannot provide this disk-eviction behavior. Eviction failures are logged and do not discard guest data. Shared/COW mappings and Sentry's own management memory are retained. This is live-process page eviction, not a checkpoint that survives host shutdown.

The eviction adapter reads the pinned gVisor version's private mapping metadata under its existing lock. It does not modify gVisor or parse diagnostic text. Each Run still uses a fresh memfd, and results remain write-sealed before host decoding. The one-shot `sandbox.Run(fn)` creates a Process, runs it and closes it automatically.

## Modules

Expand Down Expand Up @@ -81,8 +105,8 @@ err := s.Run(func() { f.OnBuild(ctx) })
| Type | Source and lifetime |
| --- | --- |
| `bind` | Host directory, served through DirectFS. Writes persist in that host directory. Host permissions still apply. |
| `tmpfs` | New Sentry filesystem for each call. Supports options such as `size`, `mode`, `uid` and `gid`; contents disappear when the call ends. |
| `proc` | Process information from this call's PID namespace. |
| `tmpfs` | New Sentry filesystem for each process. Supports options such as `size`, `mode`, `uid` and `gid`; contents persist until the process closes. |
| `proc` | Process information from the guest's PID namespace. |
| `overlay` | Combines already visible guest paths using `lowerdir` and optional `upperdir`. Writes go to the upper layer, whose filesystem determines persistence. |

An empty `Mounts` retains the default read-only host `/` and guest `/proc`. A nonempty list replaces all defaults. Its first entry must mount `bind` or `tmpfs` at `/`; subsequent mounts are applied in order, so parents and overlay layers must precede their users. Duplicate targets and unsupported types return an error. Missing directory mountpoints are prepared through Sentry's synthetic-mountpoint support, which still requires a writable parent mount. For a read-only parent, prepare the target directory beforehand or place new mountpoints under a writable tmpfs. Bind sources currently must be directories.
Expand Down
3 changes: 2 additions & 1 deletion bridge.h
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,8 @@ struct syscall_event {
typedef void (*inspect_fn)(uintptr_t, struct syscall_event *);
typedef int (*create_sentry_fn)(uintptr_t *, char *, size_t);
typedef int (*close_sentry_fn)(uintptr_t, char *, size_t);
// Each run receives a Kernel handle and startup JSON (guest, mounts and env).
// Dispatches create/run/close for a process ID within a Kernel.
// Creation supplies guest/mount/env configuration; each run supplies a fresh image fd.
typedef int (*run_sentry_fn)(uintptr_t, char *, int, uintptr_t, uintptr_t, uintptr_t,
inspect_fn, char *, size_t);

Expand Down
47 changes: 32 additions & 15 deletions guest_linux.go
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ import (
"context"
"errors"
"fmt"
"io"
"os"
"runtime"

Expand All @@ -30,23 +31,39 @@ func runGuest() (err error) {
err = fmt.Errorf("sandbox guest panicked: %v", value)
}
}()
control := os.NewFile(4, "sandbox-control")
defer control.Close()
defer unix.Close(3)
data, unmap, err := readStateImage(3, 0)
if err != nil {
return err
}
defer func() { err = errors.Join(err, unmap()) }()
var graph state.State
var fn func()
ctx := context.Background()
if _, err := graph.Load(ctx, data, &fn); err != nil {
return err
}
runtime.GC()
fn()
_, err = writeStateImage(3, int64(len(data))+8, &graph, &fn)
if err != nil {
return err
var command [1]byte
for {
if _, err := io.ReadFull(control, command[:]); err != nil {
if errors.Is(err, io.EOF) {
return nil
}
return err
}
if command[0] != 1 {
return errors.New("sandbox: invalid process command")
}
data, unmap, err := readStateImage(3, 0)
if err != nil {
return err
}
resultOffset := int64(len(data)) + 8
_, loadErr := graph.Load(context.Background(), data, &fn)
err = errors.Join(loadErr, unmap())
if err != nil {
return err
}
runtime.GC()
fn()
if _, err := writeStateImage(3, resultOffset, &graph, &fn); err != nil {
return err
}
if _, err := control.Write(command[:]); err != nil {
return err
}
}
return nil
}
207 changes: 155 additions & 52 deletions host_linux.go
Original file line number Diff line number Diff line change
Expand Up @@ -80,82 +80,108 @@ func sandboxInspect(owner C.uintptr_t, event *C.struct_syscall_event) {
}
}

// Run executes fn in a fresh guest running the same ELF. The guest enters the
// closure automatically after package initialization, without running main.
// Capture mutations are committed only after a successful guest exit.
// Captures must be exclusively owned for the duration of Run. Calls may overlap
// when their captured graphs are independent. Configuration must not change
// until all active calls return.
type processState struct {
kernel uintptr
inspection cgo.Handle
observer *inspection
graph state.State
fn func()
}

// Run executes fn in a new guest and closes it after writing captures back.
func (s *Sandbox) Run(fn func()) (err error) {
p := NewProcess(ProcessOptions{Mounts: s.Mounts, Env: s.Env, Inspect: s.Inspect})
p.owner = s
defer func() { err = errors.Join(err, p.Close()) }()
return p.Run(fn)
}

// Run executes fn in this process and pauses the guest before returning.
// Calls on one Process are serialized. Captures must be exclusively owned
// until Run returns. A failed transfer or guest execution closes the process.
func (p *Process) Run(fn func()) (err error) {
if fn == nil {
return fmt.Errorf("sandbox: nil function")
return errors.New("sandbox: nil function")
}
handleID, err := s.acquire()
if err != nil {
p.runMu.Lock()
p.mu.Lock()
if p.closed || p.failure != nil || p.owner == nil {
err = p.failure
if err == nil {
err = errors.New("sandbox: process is closed or uninitialized")
}
p.mu.Unlock()
p.runMu.Unlock()
return err
}
defer s.active.Done()
mainPC, err := guestMain()
p.runs.Add(1)
p.mu.Unlock()
transfer := false
defer func() {
if err != nil && transfer {
p.mu.Lock()
p.failure = err
p.mu.Unlock()
}
p.runs.Done()
p.runMu.Unlock()
if err != nil && transfer {
err = errors.Join(err, p.Close())
}
}()
kernelID, err := p.owner.acquire()
if err != nil {
return err
}
entryPC := reflect.ValueOf(guestEntry).Pointer()
defer p.owner.active.Done()
fd, err := newStateImage()
if err != nil {
return err
}
defer unix.Close(fd)
var graph state.State
ctx := context.Background()
resultOffset, err := writeStateImage(fd, 0, &graph, &fn)
p.native.fn = fn
resultOffset, err := writeStateImage(fd, 0, &p.native.graph, &p.native.fn)
if err != nil {
return fmt.Errorf("sandbox export: %w", err)
}
transfer = true
runtime.GC()
executable, err := os.Executable()
if err != nil {
return err
p.mu.Lock()
if p.closed {
p.mu.Unlock()
return errors.New("sandbox: process is closed")
}
mounts := s.Mounts
if len(mounts) == 0 {
mounts = []Mount{
{Type: "bind", Source: "/", Target: "/", Options: []string{"ro"}},
{Type: "proc", Target: "/proc"},
if p.native.kernel == 0 {
p.native.kernel = kernelID
p.native.observer = &inspection{fn: p.options.Inspect}
if p.options.Inspect != nil {
p.native.inspection = cgo.NewHandle(p.native.observer)
}
err = p.call("create", fd)
if err == nil {
p.owner.mu.Lock()
if p.owner.processes == nil {
p.owner.processes = make(map[uint64]*Process)
}
p.owner.processes[p.id] = p
p.owner.mu.Unlock()
}
}
env := s.Env
if env == nil {
env = os.Environ()
}
config, err := json.Marshal(struct {
Guest string `json:"guest"`
Mounts []Mount `json:"mounts"`
Env []string `json:"env"`
}{executable, mounts, env})
p.mu.Unlock()
if err != nil {
return fmt.Errorf("sandbox configuration: %w", err)
}
cConfig := C.CString(string(config))
defer C.free(unsafe.Pointer(cConfig))
i := &inspection{fn: s.Inspect}
var handle cgo.Handle
if s.Inspect != nil {
handle = cgo.NewHandle(i)
defer handle.Delete()
return err
}
var message [4096]C.char
code := C.sandbox_run(C.uintptr_t(handleID), cConfig, C.int(fd), C.uintptr_t(mainPC), C.uintptr_t(entryPC), C.uintptr_t(handle), &message[0], C.size_t(len(message)))
if code != 0 {
return fmt.Errorf("sandbox Sentry: %s", C.GoString(&message[0]))
if err := p.call("run", fd); err != nil {
return err
}
i.mu.Lock()
inspectionErr := i.err
i.mu.Unlock()
p.native.observer.mu.Lock()
inspectionErr := p.native.observer.err
p.native.observer.mu.Unlock()
if inspectionErr != nil {
return inspectionErr
}
// Seal before either decode pass, including against any fd the guest
// transferred elsewhere. Existing writable mappings make this fail.
// Each call gets a new memfd. The previous result stays sealed even if the
// resident guest retained a duplicate descriptor across calls.
if _, err := unix.FcntlInt(uintptr(fd), unix.F_ADD_SEALS, unix.F_SEAL_WRITE|unix.F_SEAL_SEAL); err != nil {
return fmt.Errorf("sandbox result seal: %w", err)
}
Expand All @@ -164,14 +190,82 @@ func (s *Sandbox) Run(fn func()) (err error) {
return fmt.Errorf("sandbox result: %w", err)
}
defer func() { err = errors.Join(err, unmap()) }()
if _, err := graph.Load(ctx, data, &fn); err != nil {
if _, err := p.native.graph.Load(context.Background(), data, &p.native.fn); err != nil {
return fmt.Errorf("sandbox import: %w", err)
}
graph = state.State{}
runtime.GC()
return nil
}

func (p *Process) call(operation string, fd int) error {
config := struct {
Operation string `json:"operation"`
Process uint64 `json:"process"`
Guest string `json:"guest,omitempty"`
Mounts []Mount `json:"mounts,omitempty"`
Env []string `json:"env,omitempty"`
}{Operation: operation, Process: p.id}
var mainPC, entryPC uintptr
if operation == "create" {
var err error
mainPC, err = guestMain()
if err != nil {
return err
}
entryPC = reflect.ValueOf(guestEntry).Pointer()
config.Guest, err = os.Executable()
if err != nil {
return err
}
config.Mounts = p.options.Mounts
if len(config.Mounts) == 0 {
config.Mounts = []Mount{{Type: "bind", Source: "/", Target: "/", Options: []string{"ro"}}, {Type: "proc", Target: "/proc"}}
}
config.Env = p.options.Env
if config.Env == nil {
config.Env = os.Environ()
}
}
data, err := json.Marshal(config)
if err != nil {
return fmt.Errorf("sandbox configuration: %w", err)
}
cConfig := C.CString(string(data))
defer C.free(unsafe.Pointer(cConfig))
var message [4096]C.char
if C.sandbox_run(C.uintptr_t(p.native.kernel), cConfig, C.int(fd), C.uintptr_t(mainPC), C.uintptr_t(entryPC), C.uintptr_t(p.native.inspection), &message[0], C.size_t(len(message))) != 0 {
return fmt.Errorf("sandbox Sentry: %s", C.GoString(&message[0]))
}
return nil
}

// Close terminates this process, including an active Run, and releases its
// resources. It is idempotent. Do not call it from this process's inspector.
func (p *Process) Close() error {
p.closeOnce.Do(func() {
p.mu.Lock()
p.closed = true
started := p.native.kernel != 0
p.mu.Unlock()
if started {
p.closeErr = p.call("close", -1)
}
p.runs.Wait()
if p.native.inspection != 0 {
p.native.inspection.Delete()
p.native.inspection = 0
}
p.native.graph = state.State{}
p.native.fn = nil
if p.owner != nil {
p.owner.mu.Lock()
delete(p.owner.processes, p.id)
p.owner.mu.Unlock()
}
})
return p.closeErr
}

func (s *Sandbox) acquire() (uintptr, error) {
s.mu.Lock()
defer s.mu.Unlock()
Expand Down Expand Up @@ -237,6 +331,15 @@ func (s *Sandbox) Close() error {
}
}
s.active.Wait()
s.mu.Lock()
processes := make([]*Process, 0, len(s.processes))
for _, p := range s.processes {
processes = append(processes, p)
}
s.mu.Unlock()
for _, p := range processes {
err = errors.Join(err, p.Close())
}
s.closeErr = err
close(s.closeDone)
return err
Expand Down
Loading
Loading