Skip to content

feat(sandbox): write agent output to the container log - #4005

Merged
krishicks merged 1 commit into
mainfrom
3928-agent-output-container-log/krishicks
Oct 1, 2026
Merged

krishicks merged 1 commit into
mainfrom
3928-agent-output-container-log/krishicks

Conversation

@krishicks

@krishicks krishicks commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Since #2726, the sandbox's main process writes its stdout and stderr into pipes that only feed the sandbox connect replay buffer, so agent output is missing from kubectl logs, docker logs, and podman logs. This PR copies that output to the container's stdout and stderr as well, restoring the pre-#2726 behavior (v0.0.110 and earlier). It also stops the Docker and VM drivers from forwarding the workload's now agent-bearing log into gateway status and events.

Related Issue

Closes #3928

Changes

  • Agent output reaches the container log. MainSession publishes each chunk it reads from the main process to the replay buffer, then copies it to the matching container stream (stdout or stderr) through a separate forwarder thread and bounded queue per stream (container_log.rs).
    • When the container runtime falls behind on a stream, that stream's reader waits instead of dropping output, so backpressure reaches the agent as it did with inherited descriptors. The other stream and attachments keep receiving output.
    • Before MainSession::finish publishes the exit, the output readers finish and queued output is drained to the container log, so an agent's final lines are not lost. A single 30-second deadline covers both. When it expires, readers waiting on the container log are released and drain the pipes into the replay buffer only, so a stalled container log cannot block exit reporting.
    • PTY-mode processes are not copied. That output carries escape sequences and echoed input, and terminal commands never reached the container log before feat(sandbox): add canonical main process #2726. exec, SSH, and SFTP sessions are not copied either.
  • Launcher log lines are unchanged. They keep their format and stay on stderr. They are now written as whole lines, and a newline is inserted first if the agent left stderr mid-line, so launcher and agent lines never merge.
  • Docker and VM drivers. Failure messages (the Ready condition and platform events republished to the sandbox event stream) no longer include the workload's output, only the supervisor's log tail. Otherwise arbitrary agent output, including anything sensitive, would land in gateway status and events: up to 80 lines of the Docker workload container's log, or the last 8 KiB of the VM guest console, which carries the launcher's stdout and stderr.
    • Docker: the supervisor's health endpoint starts only after the agent starts, so every failure path could include agent output. The supervisor-only warn! that fix(runtime): recover SSH relays and bound startup diagnostics #4011 added to the readiness failure path no longer logs the workload tail, and fix(runtime): recover SSH relays and bound startup diagnostics #4011's 1 KiB cap on status-message tails is kept.
    • VM: ProcessExited is reported whenever the VM or host supervisor exits, including after the main process ends. It keeps the host supervisor's stderr tail.
    • This matches the Podman driver, which never forwards raw workload output. The workload's output remains available through docker logs and the VM's rootfs-console.log.
  • Docs. docs/observability/accessing-logs.mdx describes where main process output appears. debug-openshell-cluster notes what the sandbox and supervisor container logs contain.

Testing

  • mise run pre-commit passes
  • Unit tests added/updated
    • container_log: output written unmodified, the newline guard for launcher lines, line assembly from fragmented writes, ordering, waiting on a stalled log without dropping output, stderr unaffected by a stalled stdout, and drain.
    • main_session: pipe output copied to the matching stream, PTY output not copied, finish waiting for queued output, finish releasing readers blocked on a saturated and permanently stalled log, stderr and attachments flowing while the stdout log is stalled, and attachments receiving a chunk whose container-log send is waiting.
    • openshell-driver-vm: process_exit_status_omits_guest_console_output drives monitor_sandbox through the VM and host-supervisor exit paths and checks that the condition and platform event omit the guest console while keeping the supervisor's stderr tail.
    • openshell-driver-docker: startup_error_log_tails_fit_grpc_header_budget (from fix(runtime): recover SSH relays and bound startup diagnostics #4011) now builds the status from the supervisor tail only.
    • Run on macOS, and on Linux (aarch64, in a container) for the Linux-only main_session tests. Linux-target clippy also passes for openshell-sandbox (with and without perf-harness) and openshell-driver-vm.
  • E2E tests added/updated (if applicable)
  • Manually verified on a local k3d gateway with an earlier revision of this change: kubectl logs <pod> -c agent showed the launcher's startup warning followed by the agent's stdout and stderr lines, and the node's container log file tagged each line with the correct stream. openshell sandbox connect still replayed from the buffer. That revision predates the waiting, PTY, and drain changes and the Docker driver change; those are covered by the unit tests above.

Checklist

  • Follows Conventional Commits
  • Commits are signed off (DCO)
  • Architecture docs updated (if applicable). Not applicable: no architecture doc describes the main-process output path.

@github-actions

Copy link
Copy Markdown

@krishicks
krishicks force-pushed the 3928-agent-output-container-log/krishicks branch from 73f0ef6 to a7aa90d Compare September 30, 2026 21:55
@devhuntr

devhuntr commented Oct 1, 2026

Copy link
Copy Markdown

I like this change, especially making the agent output available through the normal container logs again. I’m still pretty new to this kind of logging setup, but I was curious why you chose to use a separate forwarder thread and queue for each stream instead of writing directly to the container log?

@krishicks

Copy link
Copy Markdown
Collaborator Author

The channel moves the blocking write() onto a dedicated thread, so that if the container runtime falls behind reading the log, the launcher's runtime isn't blocked, and the reader's wait can be abandoned at the exit deadline so the agent's exit is still reported.

@krishicks
krishicks force-pushed the 3928-agent-output-container-log/krishicks branch from a7aa90d to 68d8a58 Compare October 1, 2026 16:40
@purp

purp commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

🤖

Potential concerns

  • VM sandboxes: agent output lands in gateway status.
    • Before: a VM sandbox creator whose VM or host supervisor exited saw a ProcessExited condition with up to 8 KiB of guest console. That console held only init-script and launcher lines.
    • With this PR: the VM guest runs openshell-sandbox launch-capability-free, which goes through the same run_boundary → delegated::spawn_workload path (main.rs:1574). Its stdout and stderr are the guest console, which is saved to rootfs-console.log.
    • Impact: arbitrary agent output, possibly sensitive, reaches the condition message and gateway status/events. That's the exposure the PR deliberately removed for Docker.
    • The new docs section also doesn't list where VM output goes.
    • Details: crates/openshell-driver-vm/src/driver.rs:4193-4220, driver.rs:113.
    • Kubernetes and Podman are unaffected. FallbackToLogsOnError is set only on the supervisor container, and neither driver reads logs.

This looks real.

The highlighted risk is that there are things in the sandbox log that shouldn't be copied out to the gateway as part of exit diagnostics; I don't know enough to say if that's an interesting risk or not.

Since #2726 the canonical main process's stdout and stderr are captured
in pipes that feed only the in-memory replay buffer used by sandbox
connect. Agent output therefore never reaches the container's own stdout
and stderr, so it is missing from kubectl logs, docker logs, and podman
logs and from anything that collects container logs. Before #2726 the
entrypoint inherited the container's descriptors and its output appeared
there.

Copy the main process's output to the launcher's stdout and stderr in
addition to the replay buffer, restoring the earlier behavior:

- Output is copied byte for byte to the matching stream from a
  forwarder thread per stream, after it is published to the replay
  buffer. When the container runtime falls behind on a stream, that
  stream's reader waits instead of dropping output, so backpressure
  reaches the agent as it did with inherited descriptors, while the
  other stream and attachments keep receiving output.
- Before the main process's exit is published, the output readers
  finish and queued output is drained to the container log, so an
  agent's final lines are not lost at shutdown. A 30 second deadline
  covers both; when it expires, readers waiting on the container log
  are released and drain the pipes into the replay buffer only, so a
  stalled container log cannot block exit reporting.
- PTY-mode processes are not copied. The terminal stream carries escape
  sequences and echoed input, and terminal commands never reached the
  container log before #2726.
- Exec, SSH, and SFTP sessions are not copied.

Launcher log lines keep their existing format and remain in the
container's stderr. They are written as whole lines, and a newline is
inserted first when the agent left stderr mid-line, so launcher and
agent lines do not merge.

The Docker and VM drivers appended the tail of the workload's output to
failure messages: Docker the workload container's log, and the VM driver
the guest console, which carries the launcher's stdout and stderr. Those
messages land in the sandbox's Ready condition and in platform events
that the gateway republishes to the sandbox event stream. With agent
output in that log, those messages would carry arbitrary agent output,
including anything sensitive the agent prints, into gateway status and
events. The supervisor starts its health endpoint only after the agent
starts, so every Docker failure path could include agent output, and the
VM driver reports one whenever the VM or host supervisor exits. Forward
only the supervisor's log tail, matching the Podman driver, which reads
the workload log solely to match fixed launcher markers and never
forwards raw workload output. The workload's output remains available
through docker logs and the VM's rootfs-console.log.

Document where main process output appears in the logging docs and the
cluster debugging skill.

Closes #3928

Signed-off-by: Kris Hicks <khicks@nvidia.com>
@krishicks
krishicks force-pushed the 3928-agent-output-container-log/krishicks branch from 68d8a58 to 049ca16 Compare October 1, 2026 17:03

@purp purp left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM 🚢

@krishicks
krishicks added this pull request to the merge queue Oct 1, 2026
Merged via the queue into main with commit 8091f66 Oct 1, 2026
94 checks passed
@krishicks
krishicks deleted the 3928-agent-output-container-log/krishicks branch October 1, 2026 17:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Sandbox agent logs written to stdout do not appear in Kubernetes container logs

3 participants