Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@
<a href="#quick-start">Quick Start</a> &middot;
<a href="#architecture">Architecture</a> &middot;
<a href="docs/ARCHITECTURE.md">Docs</a> &middot;
<a href="docs/RELEASE_NOTES.md">Release Notes</a> &middot;
<a href="#contributing">Contributing</a> &middot;
<a href="#license">License</a>
</p>
Expand Down Expand Up @@ -165,6 +166,19 @@ layer extraction, rootfs caching, networking setup, subprocess spawn, and
post-boot hooks. It returns a `*VM` handle that you use to query status, stop,
or remove the VM.

Virtio-fs ownership overrides preserve host ownership and mode. Mounts with
`OverrideUID` are prepared before startup, including read-only exports. Startup
uses bounded best-effort reporting by default: recoverable entry failures produce
one warning for the incomplete mount while safe descendants, siblings, later
mounts, and VM startup continue. Set `StrictOwnershipPreparation: true` on a
mount to fail startup before networking. The public
`virtiofs.PrepareOwnership` API is always strict and supports targeted updates
after startup. Both policies use the same descriptor-relative, symlink-confined
filesystem operations. New or changed xattrs still require host permission; no
path implicitly changes mode or ownership. See
[macOS support](docs/MACOS.md#virtio-fs-shared-directory-ownership) and
[release notes](docs/RELEASE_NOTES.md).

## Advanced Usage

For appliance-style deployments, go-microvm exposes hooks and overrides at every
Expand Down Expand Up @@ -329,6 +343,7 @@ func main() {
| `ssh` | No | ECDSA key generation and SSH client for guest communication |
| `state` | No | flock-based state persistence with atomic JSON writes |
| `rootfs` | No | Rootfs cloning with reflink (copy-on-write) support |
| `virtiofs` | No | Strict, symlink-confined `override_stat` ownership preparation for shared host trees |
| `internal/pathutil` | No | Path traversal validation for safe file operations |
| `internal/xattr` | No | Extended attribute helpers for `override_stat` ownership mapping |

Expand Down
15 changes: 12 additions & 3 deletions docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,7 +76,16 @@ in `microvm.Run()`:
config (Entrypoint, Cmd, Env, WorkingDir), applying `WithInitOverride` if
set. Writes the JSON file to `/.krun_config.json` in the rootfs.

5. **Start networking** -- Networking follows one of two paths:
5. **Prepare virtio-fs ownership** -- For every mount opting into
`OverrideUID`, including read-only exports, prepares `override_stat` metadata
with descriptor-relative, symlink-confined traversal. Startup is best-effort
by default: recoverable entry failures are bounded in one incomplete-mount
warning while safe descendants, siblings, and subsequent mounts continue.
`StrictOwnershipPreparation` instead fails before networking. The public
`virtiofs.PrepareOwnership` API remains strictly fail-fast. All mount
configuration is validated before any metadata is stamped.

6. **Start networking** -- Networking follows one of two paths:
- **Default (no `WithNetProvider`)**: Port forwards are passed to the
runner via `runner.Config`. The runner creates an in-process
VirtualNetwork (gvisor-tap-vsock) connected via a socketpair.
Expand All @@ -90,14 +99,14 @@ in `microvm.Run()`:
runs the VirtualNetwork in the caller's process and supports HTTP
services on the gateway IP.

6. **Start VM via backend** -- The `hypervisor.Backend` handles rootfs
7. **Start VM via backend** -- The `hypervisor.Backend` handles rootfs
preparation and VM launch. The default libkrun backend serializes
`runner.Config` as JSON and spawns `go-microvm-runner` as a detached
subprocess (`setsid` for new session). The runner is located by
searching: explicit path, system PATH, then next to the calling
executable. Custom backends can be provided via `WithBackend()`.

7. **Post-boot hooks** -- Runs caller-provided `PostBootHook` functions. If
8. **Post-boot hooks** -- Runs caller-provided `PostBootHook` functions. If
any hook fails, the VM is stopped and the error is returned.

### Runner Side (go-microvm-runner)
Expand Down
104 changes: 104 additions & 0 deletions docs/MACOS.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,6 +86,110 @@ uid/gid/mode to the guest. This is the same mechanism used by podman on macOS.
The xattr is set automatically during OCI layer extraction and rootfs cloning
-- no user action is needed.

### virtio-fs shared directory ownership

A `microvm.VirtioFSMount` with `OverrideUID > 0` is prepared before
networking starts, whether its export is writable or read-only. `OverrideGID`
defaults to the UID. Startup is best-effort by default: inaccessible entries,
malformed metadata, unsupported special files, and xattr errors are retained in
a bounded report and emitted as one warning for the incomplete mount. Traversal
continues through safe accessible descendants, siblings, and subsequent mounts.
Set `StrictOwnershipPreparation: true` to fail startup on the first such error.
Root/target acquisition failures and cancellation always abort startup.

```go
microvm.WithVirtioFS(
// Backward-compatible default: report incomplete preparation and continue.
microvm.VirtioFSMount{
Tag: "shared", HostPath: "/srv/vm-share", OverrideUID: 65532,
},
// Mecatl data must be complete before networking or VM startup.
microvm.VirtioFSMount{
Tag: "mecatl", HostPath: "/srv/mecatl", OverrideUID: 65532,
StrictOwnershipPreparation: true,
},
)
```

`ReadOnly` remains enforced independently by libkrun and the guest mount. It does
not skip ownership preparation or alter the backing inode's host mode:

```go
vm, err := microvm.Run(ctx, image,
microvm.WithVirtioFS(microvm.VirtioFSMount{
Tag: "shared", HostPath: "/srv/vm-share", ReadOnly: true,
OverrideUID: 65532,
}),
)
```

The same strict public API can prepare a newly created worktree within an
already exported stable root before it is registered for consumer-level guest
use, or one replaced file/subtree after a merge, without rescanning siblings or
restarting the VM:

```go
// The caller holds its normal guest/worktree synchronization here.
if err := virtiofs.PrepareOwnership(ctx, "/srv/vm-share", "worktrees/job-42", 65532, 65532); err != nil {
return err
}
registerWorktreeWithGuest("worktrees/job-42")

// After a synchronized host create or replacement:
if err := virtiofs.PrepareOwnership(ctx, "/srv/vm-share", "results/job-42", 65532, 65532); err != nil {
return err
}
```

This changes host metadata only; no dynamic mount-add API or VM restart is
implied. The running guest or virtio-fs implementation may cache attributes, so
there is no immediate cache-invalidation or visibility guarantee.

The root must be a real directory, and the selected target must be `.` or a
relative path. Trusted ancestor symlinks such as macOS `/var` are allowed, but
the final root and every explicit relative component are opened without
following symlinks. Descendant symlinks are skipped. Only directories and
regular files are supported. Keep the export root stable for the VM lifetime:
libkrun pins that host mount, so replacing the root pathname does not retarget a
running guest.

Host ownership and mode are unchanged. New metadata derives the guest mode from
the host inode. Existing metadata retains its permission, set-ID, and sticky
bits (including guest `chmod` changes), while preparation corrects the file type
and applies the requested uid/gid. Matching metadata is not rewritten. For a
sealed snapshot, guest `0600` deliberately preserves guest-owner readability
while host `0400` narrows the backing inode; prepare while it is `0600`, then
explicitly narrow it:

```go
if err := virtiofs.PrepareOwnership(ctx, root, "snapshot", 65532, 65532); err != nil {
return err
}
if err := os.Chmod(filepath.Join(root, "snapshot"), 0o400); err != nil {
return err
}
```

A later matching preparation only reads the metadata, so it does not need to
rewrite the xattr. This is an explicit caller operation; preparation never calls
`chmod`, `chown`, or widens permissions. Host mode `0400` is not itself a
workaround for guest writes—use a read-only export for enforcement.

Creating or changing metadata requires the host OS permission to write xattrs.
An ordinary unprivileged user therefore cannot normally annotate an unannotated
`0400` file. `PrepareOwnership` returns a path-specific permission error and
leaves its host mode and IDs intact. A read-only virtio-fs export does not grant
xattr-write permission on its backing inodes.

Preparation is nontransactional. Callers must synchronize it with host rename,
creation, and replacement and with guest access or `chmod`. There is no atomic
visibility or cache-invalidation guarantee; a descriptor held across replacement
continues to refer to the old inode. A hard link in the authorized tree
authorizes changing the xattr on that inode, including names outside the tree.
The caller must also trust and protect the root's parent while the root descriptor
is acquired; subsequent traversal is descriptor-relative and confined beneath
the acquired root.

## Guest Networking

On macOS, libkrun's Hypervisor.framework backend pre-configures the guest
Expand Down
11 changes: 11 additions & 0 deletions docs/RELEASE_NOTES.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
# Release notes

## v0.0.41

### Virtio-fs ownership preparation

- Added public `virtiofs.PrepareOwnership(ctx, root, relativePath, uid, gid)` for strict full-tree or targeted `user.containers.override_stat` preparation on macOS and Linux.
- Mounts with `OverrideUID`, including read-only exports, prepare ownership before networking. Startup is best-effort by default: recoverable entry failures produce one bounded incomplete-mount warning while safe descendants, siblings, later mounts, and startup continue.
- Added per-mount `StrictOwnershipPreparation` to abort before networking and VM startup on the first preparation failure. The public API remains strict, and cancellation or mount root/target acquisition failures always abort.
- Both policies share descriptor-relative, `O_NOFOLLOW` traversal. Explicit symlinks fail, descendant symlinks are skipped, and host ownership, mode, and export flags are unchanged.
- Existing override mode bits and `OverrideGID` defaulting are preserved. Mounts without `OverrideUID` remain unmodified, and Linux user-namespace behavior is unchanged. New or changed metadata still requires host xattr-write permission; matching metadata is not rewritten.
46 changes: 35 additions & 11 deletions docs/SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -464,17 +464,41 @@ capsh --addamb=cap_chown -- -c '/path/to/your-binary'

### override_stat xattr (macOS and Linux)

go-microvm also sets the `user.containers.override_stat` extended attribute on
extracted files so that libkrun's virtiofs server reports correct ownership to
the guest. This is the same mechanism that podman uses on macOS.

On Linux, the xattr is set on regular files and directories. The kernel
restricts `user.*` xattrs on symlinks and special files, so those are silently
skipped. Once libkrun's Linux virtiofs passthrough adds support for reading
these xattrs (the same support already exists on macOS), file ownership in the
guest will be correct without requiring `CAP_CHOWN`.

See the `internal/xattr` package for details.
go-microvm sets `user.containers.override_stat` on extracted files so libkrun's
virtiofs server reports intended guest ownership without changing host uid, gid,
or mode. This is the mechanism podman uses on macOS. Image extraction and hooks
retain their single-entry best-effort behavior for compatibility.

For shared virtio-fs trees, mounts with `OverrideUID > 0`, including read-only
exports, use the same descriptor-relative ownership walker before networking.
Startup is best-effort by default: recoverable per-entry failures are retained in
a bounded report, emitted as one warning for the incomplete mount, and do not
stop safe descendants, siblings, later mounts, or VM startup. Setting
`StrictOwnershipPreparation` on a mount makes the first failure fatal. Root or
target acquisition failures and cancellation are always fatal. The public
`virtiofs.PrepareOwnership` API is always strict.

Matching existing metadata is read without a rewrite, but new or changed
metadata requires OS permission to write xattrs. Thus an unprivileged caller can
receive a path-specific permission error for an unannotated `0400` backing file,
while host mode and IDs remain unchanged. New metadata derives guest mode from
the host inode; later preparation preserves guest mode bits recorded in
`override_stat`. Read-only export enforcement is independent and does not skip
preparation or change the backing inode. The walker uses descriptor-relative,
`O_NOFOLLOW` traversal after opening the authorized root. Explicit root/target
symlinks fail, descendant symlinks are skipped, and regular files/directories
are the only supported types. This prevents mutable intermediate symlinks from
redirecting the walk, but does not make preparation transactional. The caller
must trust the root's parent during root descriptor acquisition and synchronize
rename, creation, replacement, and guest chmod. Hard links authorize their
shared inode, including names outside the tree. There are no cache invalidation
or atomic guest-visibility guarantees.

Public strict errors and startup incomplete reports are path-specific. Valid
matching metadata is not rewritten. Neither policy changes host ownership or
mode, invokes `chmod`/`chown`, widens permissions, watches the tree, or claims
atomicity. On Linux, this does not alter user-namespace behavior; it only
prepares metadata for libkrun versions that consume `override_stat`.

## File Permissions

Expand Down
14 changes: 14 additions & 0 deletions internal/xattr/noattr_darwin.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
// SPDX-FileCopyrightText: Copyright 2025 Stacklok, Inc.
// SPDX-License-Identifier: Apache-2.0

//go:build darwin

package xattr

import (
"errors"

"golang.org/x/sys/unix"
)

func isNoAttribute(err error) bool { return errors.Is(err, unix.ENOATTR) }
14 changes: 14 additions & 0 deletions internal/xattr/noattr_linux.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
// SPDX-FileCopyrightText: Copyright 2025 Stacklok, Inc.
// SPDX-License-Identifier: Apache-2.0

//go:build linux

package xattr

import (
"errors"

"golang.org/x/sys/unix"
)

func isNoAttribute(err error) bool { return errors.Is(err, unix.ENODATA) }
40 changes: 40 additions & 0 deletions internal/xattr/preparation.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
// SPDX-FileCopyrightText: Copyright 2025 Stacklok, Inc.
// SPDX-License-Identifier: Apache-2.0

package xattr

import (
"fmt"
"strings"
)

const maxPreparationErrors = 16

// PreparationReport describes recoverable entries that could not be prepared.
// Error details are bounded so a large tree cannot produce an unbounded report.
type PreparationReport struct {
Errors []error
Omitted int
}

// Complete reports whether every selected entry was prepared.
func (r PreparationReport) Complete() bool { return len(r.Errors) == 0 && r.Omitted == 0 }

func (r PreparationReport) Error() string {
parts := make([]string, 0, len(r.Errors)+1)
for _, err := range r.Errors {
parts = append(parts, err.Error())
}
if r.Omitted > 0 {
parts = append(parts, fmt.Sprintf("%d additional errors omitted", r.Omitted))
}
return strings.Join(parts, "; ")
}

func (r *PreparationReport) add(err error) {
if len(r.Errors) < maxPreparationErrors {
r.Errors = append(r.Errors, err)
} else {
r.Omitted++
}
}
Loading
Loading